When We Sound Alike: Lexical and Acoustic Alignment in Casual Conversation
- Published
- Thu, Oct 01, 2026
- Tags
- rotm
- Contact

Conversational speech is marked by alignment, which is the tendency of speakers to become more similar to each other over time, in their word choice, their prosody, and the timing of their turns. While prior work has studied lexical and acoustic-prosodic alignment largely in isolation, it remains unclear how these two dimensions interact within a single conversational exchange.
In this paper, we investigate how lexical and semantic similarity relate to acoustic-prosodic alignment on a turn-by-turn basis, using selected turn pairs from a corpus of casual Austrian German conversations. Using Conditional Random Forests and Conditional Inference Trees, we find that higher lexical similarity between turns is most strongly predicted by pronunciation matching and articulation-rate alignment, while higher semantic similarity co-occurs with pitch-minimum alignment and, in longer turns, with articulation-rate alignment as well. The figure shows this for lexical similarity on the full dataset: the top panel ranks acoustic-prosodic features by their contribution to predicting lexical similarity, and the bottom panel shows how the two strongest predictors, which are the difference in pronunciation distance and articulation-rate difference, split the data into groups with systematically different similarity scores. Speaker pairs who align closely in both pronunciation distance and speaking rate show markedly higher lexical similarity than pairs who diverge on either dimension. These findings point to a systematic link between what speakers say and how they say it, and may help inform the design of more naturally adaptive dialogue systems.
The results are published and won the Best Student Paper Award at 29th International Conference on Text, Speech and Dialogue. Congratulations!
Browse the Results of the Month archive.
