AGENTIC-SYSTEMSApril 3, 20261 min read13 views

The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation

Researchers have introduced RiDiC, a configurable pipeline for generating multilingual datasets to evaluate the factuality of LLMs in long-form generation, revealing that even frontier models struggle with hallucinations on less popular topics.

Computer Science > Computation and Language

Title:The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation

View PDF HTML (experimental)Abstract:We present a configurable pipeline for generating multilingual sets of entities with specified characteristics, such as domain, geographical location and popularity, using data from Wikipedia and Wikidata. These datasets are intended for evaluating the factuality of LLMs' long-form generation, thereby complementing evaluation based on short-form QA datasets. We present the RiDiC dataset as an example of this approach. RiDiC contains 3,000 entities from three domains -- rivers, natural disasters, and car models -- spanning different popularity tiers. Each entity is accompanied by its geographical location, English and Chinese names (if available) and relevant English and Chinese Wikipedia content, which is used to evaluate LLMs' responses. Generations about RiDiC entities were obtained from three LLMs in English and Chinese. These were then evaluated using a third-party factuality checker, which showed that entities from our dataset caused even frontier models to hallucinate. To facilitate the evaluation of LLMs' long-form factuality in multiple languages, the code, data, and generation/evaluation scripts have been released.

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Source: arXiv cs.AI Recent

More in this category

agentic-systems

VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification

Researchers have introduced VeriSimpl, a novel framework that leverages LLMs and simplification-based verification to accurately translate natural language descriptions into executable optimization formulations.

agentic-systems

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

Researchers introduced SonicSampler, a unified suite of tile-aware Triton kernels designed to accelerate LLM sampling and speculative verification. By fusing the entire pipeline into a single CUDA Graph-compatible kernel, SonicSampler achieves up to 16x speedups over state-of-the-art baselines.

agentic-systems

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Researchers have introduced PersonaTrail, a novel benchmark evaluating personalized web agents through realistic browsing histories. Alongside it, the PACMem memory framework structures raw browsing data to significantly boost agent performance in inferring user preferences.

agentic-systems

Benchmarking the Personalization Capabilities of Large Language Models

A new study introduces SDR-Bench and SDR-Arena to benchmark the personalization capabilities of LLMs in two-party persuasion scenarios, revealing a personalization plateau among frontier models.

agentic-systems

Incomplete Prompt Jailbreaks in Large Language Models

Researchers have conceptualized Incomplete Prompt Jailbreaks (IPJ), a vulnerability where incomplete harmful prompts bypass LLM safeguards due to delayed refusal behaviors. The discovery of functional neurons offers a novel path for precise neuron-level defense interventions.

agentic-systems

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

A new benchmark evaluates five leading large language models on multi-sensor physical hazard assessments, revealing a critical vulnerability when evaluating combined sensor risks below individual thresholds. Despite near-perfect accuracy on single-sensor breaches, all models failed to issue precautionary signals for multi-sensor hazards.