Skip to main navigation Skip to search Skip to main content

LAW OF THE WEAKEST LINK: CROSS CAPABILITIES OF LARGE LANGUAGE MODELS

  • Ming Zhong
  • , Aston Zhang
  • , Xuewei Wang
  • , Rui Hou
  • , Wenhan Xiong
  • , Chenguang Zhu
  • , Zhengxing Chen
  • , Liang Tan
  • , Chloe Bi
  • , Mike Lewis
  • , Sravya Popuri
  • , Sharan Narang
  • , Melanie Kambadur
  • , Dhruv Mahajan
  • , Sergey Edunov
  • , Jiawei Han
  • , Laurens van der Maaten

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term cross capabilities. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce CROSSEVAL, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that current LLMs consistently exhibit the “Law of the Weakest Link,” where cross-capability performance is significantly constrained by the weakest component. Across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight LLMs' underperformance in cross-capability tasks, emphasizing the need to identify and improve their weakest capabilities as a key research priority. The code, benchmarks, and evaluations are available on our project website.

Original languageEnglish (US)
Title of host publication13th International Conference on Learning Representations, ICLR 2025
PublisherInternational Conference on Learning Representations, ICLR
Pages82214-82277
Number of pages64
ISBN (Electronic)9798331320850
StatePublished - 2025
Event13th International Conference on Learning Representations, ICLR 2025 - Singapore, Singapore
Duration: Apr 24 2025Apr 28 2025

Publication series

Name13th International Conference on Learning Representations, ICLR 2025

Conference

Conference13th International Conference on Learning Representations, ICLR 2025
Country/TerritorySingapore
CitySingapore
Period4/24/254/28/25

ASJC Scopus subject areas

  • Language and Linguistics
  • Computer Science Applications
  • Education
  • Linguistics and Language

Fingerprint

Dive into the research topics of 'LAW OF THE WEAKEST LINK: CROSS CAPABILITIES OF LARGE LANGUAGE MODELS'. Together they form a unique fingerprint.

Cite this