AGENTIC-SYSTEMSMarch 23, 20261 min read18 views

ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models

Researchers introduce ItinBench, a new benchmark that evaluates LLMs' planning capabilities by integrating spatial and verbal reasoning. The findings reveal that top models still struggle to maintain consistent performance across multiple cognitive dimensions.

Computer Science > Artificial Intelligence

Title:ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models

View PDF HTML (experimental)Abstract:Large language models (LLMs) with advanced cognitive capabilities are emerging as agents for various reasoning and planning tasks. Traditional evaluations often focus on specific reasoning or planning questions within controlled environments. Recent studies have explored travel planning as a medium to integrate various verbal reasoning tasks into real-world contexts. However, reasoning tasks extend beyond verbal reasoning alone, and a comprehensive evaluation of LLMs requires a testbed that incorporates tasks from multiple cognitive domains. To address this gap, we introduce ItinBench, a benchmark that features one task of spatial reasoning, i.e., route optimization, into trip itinerary planning while keeping the traditional verbal reasoning tasks. ItinBench evaluates various LLMs across diverse tasks simultaneously, including Llama 3.1 8B, Mistral Large, Gemini 1.5 Pro, and GPT family. Our findings reveal that LLMs struggle to maintain high and consistent performance when concurrently handling multiple cognitive dimensions. By incorporating tasks from distinct human-level cognitive domains, ItinBench provides new insights into building more comprehensive reasoning testbeds that better reflect real-world challenges. The code and dataset: this https URL

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Source: arXiv cs.AI Recent

More in this category

agentic-systems

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google has introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, designed to deliver higher token efficiency, lower latency, and enhanced capability for scaling agentic AI workflows.

agentic-systems

Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models

Researchers introduce Generative Ontology Induction (GOI), a domain-agnostic framework that automatically extracts structured ontologies from document corpora using LLMs. Achieving 95-100% structural coverage, GOI addresses a major bottleneck in knowledge-intensive AI systems.

agentic-systems

JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models

Researchers have proposed JUMP, a novel single-pass membership inference attack designed for fine-tuned discrete diffusion language models (dLLMs). By leveraging the unique properties of dLLMs, JUMP significantly improves detection accuracy while drastically reducing the number of required queries.

agentic-systems

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL

Researchers demonstrate that Masked Diffusion Language Models (MDLMs) serve as highly effective, steerable text-based world models for agentic reinforcement learning. By leveraging bidirectional denoising, MDLMs outperform autoregressive models four times their size in coherence, groundedness, and rollout diversity.

NOW LET US Related – Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment

agentic-systems

Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment

A new study demonstrates that small language models (SLMs) under 3 billion parameters can serve as highly capable local experts for specialized tasks. By combining structured benchmarking with low-cost parameter-efficient fine-tuning (PEFT), institutions can achieve AI autonomy without relying on expensive hardware.

agentic-systems

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection

Researchers have introduced PlanFlip, a novel prompt injection attack framework targeting the planning phase of multi-agent LLM systems. The study reveals critical security blind spots in homogeneous agent pipelines and demonstrates that reasoning-augmented models like DeepSeek-R1 exhibit strong resistance.