An in-the-wild benchmark for AI agents in the production harness.
-
Updated
Aug 17, 2026 - Python
An in-the-wild benchmark for AI agents in the production harness.
Gen AI Evaluation Toolkit on AWS is a flexible, cloud-native accelerator built on AWS serverless architecture that enables comprehensive evaluation of generative AI applications.
DataClawEval: A Benchmark for Engineering Data Agents in Real Industrial Harness
A SnitchBench-style benchmark inverted for the dark-forest problem: does a listener AI alert humans about an alien signal when alerting may doom humanity?
🚀 AI Evolution Factory - From evaluation tool to continuous AI self-improvement platform. Agentic evaluation, auto-finetuning, global P2P testing, and hardware telemetry for local LLMs
A small benchmark for agent skills, verification artifacts, and fresh-session resumability.
Pipeline to investigate structured reasoning and instruction adherence in multimodal LLMs
Runnable lab measuring invalidation/staleness in agent memory — paper: Are We Ready For An Agent-Native Memory System? (arXiv:2606.24775)
Deterministic synthetic fixtures and hard-gate scoring for agentic biosafety safeguard routing.
Two-tier honest-evaluation harness for agentic RTL design (research, WIP)
Proof-bound evaluator stress testing with oracle-witnessed reward-hacking exploits and replayable evidence.
Measure AI agents’ performance with standardized tests across 314 tasks, 33 domains, and 4 difficulty levels for clear, reproducible comparison.
A Harbor-format RL environment for pass@k estimation, with a verifier built to resist reward hacking. 17 tests, 10 adversarial solutions rejected, amd64 CI.
A GPU-backed Harbor/Terminal-Bench RL environment for data-poisoning defense: containerized task, oracle solution, and a 25-check verifier that rejects 17 reward-hacking solutions.
To associate your repository with the agentic-evaluation topic, visit your repo's landing page and select "manage topics."