Skip to main content
Science AI

BenchShield preprint targets reward hacking in AI-agent benchmarks

A new preprint describes BenchShield, an instrumentation layer for detecting reward hacking in interactive LLM-agent benchmarks. The system uses a finite lifecycle model of reward-relevant events, combining static taint analysis before a run with runtime evidence to attribute agent activity. Its authors evaluated 456 labeled trajectories drawn from more than 31,000 public runs across three benchmarks. They report higher recall and same-vector coverage than an agentic scanner baseline, lower per-task cost, and 96% runtime detection accuracy. These are preprint results rather than an independently validated deployment.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Capital
AI Capital

Epsilon Health exits stealth with $20M Series A led by AlleyCorp

Epsilon Health, an AI-enabled radiology practice that contracts with radiologists using its software to generate image reports faster, emerged from stealth with a $20 million Series A round led by AlleyCorp, CEO Rustin Rassoli tol...

Science Infrastructure
Science Infrastructure

LSU demonstrates room-temperature multiphoton quantum reservoir

Louisiana State University researchers reported a room-temperature optical platform that uses bright classical light and photon-number-resolving measurements to access multiphoton quantum behavior for information processing. The ...

Security AI
Security AI

Vendor study finds AI-related SOC alerts up 685% but almost all noise

A security vendor's review of roughly 16.9 million enterprise SOC alerts found about 73,000, or 0.43%, were tied to AI tools and agents, with that volume up 685% between February and June 2026. The vendor sorted the AI-related al...

AI Products
AI Products

Specific releases Real-SWE benchmark for private enterprise codebases

Specific released Real-SWE, a benchmark for frontier AI coding agents using tasks from licensed private production codebases. The company says the tasks cover workflows such as billing, tax calculation and customer migrations, and...