TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
Abstract
Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challenge persists in part because existing group conversation datasets are often limited to short-term lab settings with contrived tasks, failing to capture the long-term social dynamics of real-world teams. To bridge this gap, we introduce TIDES, a high-resolution longitudinal dataset tracking 12 university project teams over a full semester. Comprising 76,434 utterances in both English and Korean from in-person meetings, TIDES provides a naturalistic record of teams working on self-managed projects. Our socio-structural annotations-covering interaction types, emergent roles, and development stages-allow for modeling of team evolution over months. Experiments demonstrate that models fine-tuned on TIDES show significant improvements in next-speaker prediction (64.53%, outperforming proprietary models in zero-shot) and, through cross-domain transfer, achieve comparable performance (within 2.1 percentage points) to published state-of-the-art on the AMI Meeting Corpus with approximately 42% less training data. However, preference ratings from human evaluators reveal an orthogonality between structure and content: while socio-structural data enhances next-turn prediction, it does not necessarily yield more human-preferred utterances. This suggests that structural and semantic modeling in group dynamics are more decoupled than previously assumed.