Abstract
Are distributional learning mechanisms capable of complex linguistic inferences requiring compositional generalisation? This question has become contentious with the development of large language models, which mimic human language abilities in many ways, but which struggle with compositional generalisation. We investigated a set of qualitatively different distributional models (word co-occurrence models, graphical models, recurrent neural networks, and Transformers), by training them on a carefully controlled artificial language containing combinatorial dependencies involving multiple words, and then testing them on novel sequences containing distributionally overlapping combinatorial dependencies. In this work, we show that graphical network models and Transformers, but not co-occurrence space models and recurrent neural networks, were able to perform compositional generalisation. This work demonstrates that the kinds of distributional models that can perform compositional generalisation are those that can represent words both individually and as a part of the phrases in which they participate.
| Original language | English (US) |
|---|---|
| Pages (from-to) | 413-441 |
| Number of pages | 29 |
| Journal | Language, Cognition and Neuroscience |
| Volume | 40 |
| Issue number | 3 |
| DOIs | |
| State | Published - 2025 |
Keywords
- Distributional semantics
- compositionality
- language model
- neural network
- semantic plausibility
ASJC Scopus subject areas
- Language and Linguistics
- Experimental and Cognitive Psychology
- Linguistics and Language
- Cognitive Neuroscience
Fingerprint
Dive into the research topics of 'Success and failure of compositional generalisation in distributional models of language'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS