Attention & context scaling benchmark for ElizaOS agents. Measures whether an agent selects the correct action and context as cognitive load increases.
Signature output: an attention scaling curve — accuracy plotted against context load.
cd packages/benchmarks/adhdbench
pip install -e .
# List scenarios
python scripts/run_benchmark.py list
# Compute baselines (no LLM needed)
python scripts/run_benchmark.py baselines
# Quick run (L0 only, 2 scale points, ~5 min)
python scripts/run_benchmark.py run --quick --model openai/gpt-oss-120b --provider openai
# Full run (all levels, all scales, both configs)
python scripts/run_benchmark.py run --full --model gpt-4o --provider openai--provider defaults to eliza (the ElizaOS TypeScript benchmark bridge
against the real AgentRuntime). Other choices: mock-passthrough (deterministic
smoke), openai, cerebras, groq, openrouter, vllm.
45 authored scenarios across 3 levels (a further 450 generated edge
variants are available with --expand-scenarios, 495 total):
| Level | Count | Tests |
|---|---|---|
| L0: Action Dispatch | 20 | Single-turn: does the agent pick the right action? |
| L1: Context Tracking | 15 | Multi-turn: buried instructions, entity tracking, distraction resistance |
| L2: Complex Execution | 10 | Multi-step tasks: add contact -> send message -> schedule follow-up |
5 scale points from 10 to 200 registered actions.
50 distractor actions across 9 domains (DeFi, social, productivity, files, communication, analytics, moderation, content, gaming) that create semantic disambiguation pressure against bootstrap actions.
2 configurations: basic (no advanced features) vs full (advancedMemory + advancedPlanning).
2 baselines: random (~30%) and always-REPLY (~49%) for score calibration.
Deterministic, binary. 7 outcome types:
ACTION_MATCH/ACTION_NOT_MATCH— correct action selected (or forbidden action avoided)TEXT_CONTAINS/TEXT_NOT_CONTAINS— response content checksPARAM_MATCH— action parameters present in responseMEMORY_RECALLED— fact from earlier conversation appears in responsePROVIDERS_REQUESTED— specific context providers were invoked
- Scaling curve (console + markdown)
- Per-scenario scores with turn-level detail
- JSON traces for debugging
- Markdown report with failure analysis
elizaos_adhdbench/
types.py frozen scenario/result types + default scale points
config.py all tuneable axes
scenarios.py 45 authored scenarios + edge-variant expansion
distractor_plugin.py 50 distractor actions across 9 domains
evaluator.py 7 deterministic evaluators + scoring
baselines.py random + always-REPLY baselines
runner.py orchestration loop (mock-passthrough path)
openai_runner.py OpenAI-compatible provider runner (openai/cerebras/groq/openrouter/vllm)
reporting.py markdown, JSON, ASCII curves
tests/ pytest suite (150 tests)
scripts/run_benchmark.py CLI with run, baselines, list commands
pip install -e ".[dev]"
pytest tests/ -v