

ACL 2026 BoF : What Makes an Enterprise Agent Benchmark Useful? Tasks, Tools & Trust
Note: Open to only ACL Main Conference Attendees
LLM agents are moving from chat into enterprise workflows: industrial asset operations, financial analysis, customer support, scientific discovery. There, evaluation stops being a leaderboard exercise and becomes a question of safety, cost, and trust.
Yet most public agent benchmarks still measure protocol fluency or single-tool calls in toy domains, leaving the hard questions open. Does the agent retrieve the right tool? Plan the right sequence? Recover from failure? Hallucinate when it shouldn't? Cost what its owner can afford?
This BoF brings together researchers and practitioners who build, evaluate, and deploy enterprise-grade agents to share lessons learned and identify common open problems. We will draw on the experience of AssetOpsBench (NeurIPS 2025; AAAI 2026 Lab), an industrial asset-operations benchmark with 450+ scenarios, 1.9k+ GitHub stars, and a CODS 2025 community challenge that drew 300+ submissions across 149 teams, to discuss what worked, what didn't, and what's still missing.
Discussion topics:
Task design. What separates an enterprise-realistic task from a toy task? Multi-step workflows, ambiguous instructions, domain knowledge.
Evaluation beyond pass@1. Trajectory-level metrics, hallucination taxonomies, execution quality, cost-aware scoring. Are leaderboards even the right format for enterprise agents?
Tool ecosystems and MCP. Evaluating agents that must discover and orchestrate tools, not just call them, across heterogeneous enterprise environments.
Privacy-aware online evaluation. Running competitions where private data, hidden test sets, and reproducibility must coexist.
Community challenges as benchmark validation. What 300+ submissions to AssetOpsBench-Live taught us about which design patterns actually transfer.
Open problems. Multi-agent failure modes, world-model evaluation, domain transfer, and the gap between research benchmarks and production deployment.
Format: brief framing remarks (~10 min), then ⚡ short lightning talks, then 🗣️ open discussion organized around participant-proposed topics. Researchers, practitioners, and anyone building or breaking enterprise agents are welcome.
Bring open questions, war stories, or both.
🎤 Call for Lightning Talks (≤5 min)
We're looking for 5-6 lightning speakers. Have a benchmark, an evaluation method, a failure mode, or a strong opinion on enterprise agent eval? Pitch it.
5 minutes, 1 to 3 slides (or none)
Topics: hard task design, eval beyond pass@1, MCP/tool-orchestration eval, privacy-aware online eval, "benchmarks that fooled us"
Submit a 1 to 2-sentence pitch plus name/affiliation: Contact the host
Pitch by July 4. Lineup confirmed July 5.