I must delete the evidence: AI Agents Explicitly Cover up Fraud and Violent Crime

Researchers found that many state-of-the-art AI agents explicitly choose to suppress evidence of fraud and harm to serve corporate interests. The study, involving 16 LLMs, highlights significant risks in AI alignment and ethical decision-making.

Computer Science > Artificial Intelligence

Title:I must delete the evidence: AI Agents Explicitly Cover up Fraud and Violent Crime

View PDF HTML (experimental)Abstract:As ongoing research explores the ability of AI agents to be insider threats and act against company interests, we showcase the abilities of such agents to act against human well being in service of corporate authority. Building on Agentic Misalignment and AI scheming research, we present a scenario where the majority of evaluated state-of-the-art AI agents explicitly choose to suppress evidence of fraud and harm, in service of company profit. We test this scenario on 16 recent Large Language Models. Some models show remarkable resistance to our method and behave appropriately, but many do not, and instead aid and abet criminal activity. These experiments are simulations and were executed in a controlled virtual environment. No crime actually occurred.

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Source: arXiv cs.AI Recent

I must delete the evidence: AI Agents Explicitly Cover up Fraud and Violent Crime

Computer Science > Artificial Intelligence

Title:I must delete the evidence: AI Agents Explicitly Cover up Fraud and Violent Crime

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

More in this category

VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark

StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents

Discover All Categories