A Cross-Linguistic Pressure for Uniform Information Density in Word Order (bibtex)
by Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn, Ryan Cotterell, Richard Futrell, Roger Levy
Abstract:
While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1
Reference:
A Cross-Linguistic Pressure for Uniform Information Density in Word OrderThomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn, Ryan Cotterell, Richard Futrell, Roger LevyTransactions of the Association for Computational Linguistics, 2023.
Bibtex Entry:
@article{clark2023pressure,
  author    = {Clark, Thomas Hikaru and Clara Meister and Tiago Pimentel and Michael Hahn and Ryan Cotterell and Richard Futrell and Roger Levy},
    title = "{A Cross-Linguistic Pressure for Uniform Information Density in Word Order}",
    journal = {Transactions of the Association for Computational Linguistics},
    volume = {11},
    pages = {1048-1065},
    year = {2023},
    month = {08},
    abstract = "{While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the effort to identify these pressures, prior work has compared real and counterfactual word orders. Yet one functional pressure has been overlooked in such investigations: The uniform information density (UID) hypothesis, which holds that information should be spread evenly throughout an utterance. Here, we ask whether a pressure for UID may have influenced word order patterns cross-linguistically. To this end, we use computational models to test whether real orders lead to greater information uniformity than counterfactual orders. In our empirical study of 10 typologically diverse languages, we find that: (i) among SVO languages, real word orders consistently have greater uniformity than reverse word orders, and (ii) only linguistically implausible counterfactual orders consistently exceed the uniformity of real orders. These findings are compatible with a pressure for information uniformity in the development and usage of natural languages.1}",
    issn = {2307-387X},
    doi = {10.1162/tacl_a_00589},
    url = {https://doi.org/10.1162/tacl\_a\_00589},
    eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/tacl\_a\_00589/2154495/tacl\_a\_00589.pdf},
}
Powered by bibtexbrowser