AI Benchmark & Evaluation Engineer · Bengaluru, India
I build the tasks that break AI agents.
I design terminal-based agentic benchmarks in the Harbor / Terminal-Bench format, calibrated so frontier models fail more often than they pass. B.E in Computer Science and Engineering and M.Tech in Cybersecurity @ VTU.
- 280+ Terminal-Bench tasks authored, ~270 accepted — Docker-isolated environments, deterministic test oracles, canary-string hygiene
- Difficulty calibration against frontier models: iterative hardening until pass rates land at ≤3/5 runs (medium/hard thresholds)
- Anti-pattern-matching design — tasks built so models must actually reason through the environment instead of recognizing a shape they've seen before
- Milestone and non-milestone task formats, multi-step verification, reproducible harnesses
- 500+ completed evaluation tasks across Snorkel AI, Appen, Toloka, Mindrift, AfterQuery, CrowdGen — consistently strong quality ratings
- Code and PR evaluation workflows, IDE-arena dataset creation, model output rating and ranking
- Agent-config authoring (
CLAUDE.md) for reproducible, repeatable task-creation pipelines
Python Bash TypeScript Docker pytest Linux Git
Harbor Terminal-Bench MCP LLM evaluation tmux asciinema
- Benchmarks that resist pattern-matching rather than reward it
- Deterministic, reproducible eval harnesses — no flaky oracles
- Honest difficulty calibration: a task is only "hard" if the numbers say so