AGENTIC-SYSTEMSApril 3, 20261 min read17 views

Brevity Constraints Reverse Performance Hierarchies in Language Models

New research identifies that larger language models often underperform smaller ones due to scale-dependent verbosity that introduces errors. By applying brevity constraints, researchers successfully reversed this performance hierarchy, improving large model accuracy by 26 percentage points.

Computer Science > Computation and Language

Title: Brevity Constraints Reverse Performance Hierarchies in Language Models

Standard evaluation protocols reveal a counterintuitive phenomenon: on 7.7% of benchmark problems spanning five datasets, larger language models underperform smaller ones by 28.4 percentage points despite 10-100x more parameters. Through systematic evaluation of 31 models (0.5B-405B parameters) across 1,485 problems, we identify the mechanism as spontaneous scale-dependent verbosity that introduces errors through overelaboration.

Causal intervention experiments demonstrate this reflects correctable prompt design rather than fundamental capability limitations. Constraining large models to produce brief responses improves accuracy by 26 percentage points and reduces performance gaps by up to two-thirds. Most critically, brevity constraints completely reverse performance hierarchies on mathematical reasoning and scientific knowledge benchmarks, with large models achieving 7.7-15.9 percentage point advantages over small models -- direct inversions of the original gaps.

These reversals prove large models possess superior latent capabilities that universal prompting masks. We validate findings through three independent contamination tests and demonstrate inverse scaling operates continuously across the full parameter spectrum, with dataset-specific optimal scales ranging from 0.5B to 3.0B parameters. Our results establish that maximizing large model performance requires scale-aware prompt engineering rather than universal evaluation protocols, with immediate implications for deployment: prompt adaptation simultaneously improves accuracy and reduces computational costs.

Source: arXiv cs.AI Recent

More in this category

agentic-systems

VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification

Researchers have introduced VeriSimpl, a novel framework that leverages LLMs and simplification-based verification to accurately translate natural language descriptions into executable optimization formulations.

agentic-systems

Semi-Supervised Text-Attributed Graph Distillation

Researchers propose a unified semi-supervised framework guided by Wasserstein Distance to tackle scalability bottlenecks in Text-Attributed Graphs (TAGs). The approach achieves a state-of-the-art trade-off between performance and data compression for both GNN- and LLM-based downstream tasks.

agentic-systems

Benchmarking the Personalization Capabilities of Large Language Models

A new study introduces SDR-Bench and SDR-Arena to benchmark the personalization capabilities of LLMs in two-party persuasion scenarios, revealing a personalization plateau among frontier models.

agentic-systems

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

Researchers introduced SonicSampler, a unified suite of tile-aware Triton kernels designed to accelerate LLM sampling and speculative verification. By fusing the entire pipeline into a single CUDA Graph-compatible kernel, SonicSampler achieves up to 16x speedups over state-of-the-art baselines.

agentic-systems

Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference

A new research study introduces a constrained multi-source inference framework to accurately reconstruct power distribution system topology. Validated on over 8,000 smart meters in the U.S., the method achieves over 95% accuracy while drastically reducing computational overhead.

agentic-systems

Incomplete Prompt Jailbreaks in Large Language Models

Researchers have conceptualized Incomplete Prompt Jailbreaks (IPJ), a vulnerability where incomplete harmful prompts bypass LLM safeguards due to delayed refusal behaviors. The discovery of functional neurons offers a novel path for precise neuron-level defense interventions.