# InfraBench > InfraBench is a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and operational lifecycle, with fine-grained risk assessment. Published at HotInfra '26 (co-located with ISCA '26). Twelve seed tasks span hardware (L1), local systems (L2), distributed systems (L3), and user applications (L4). Scoring goes beyond binary pass/fail to cover durability, invariants, cleanup, and risk. Affiliations: University of Wisconsin–Madison and Iowa State University. ## Docs - [Project website](https://infraben.ch/): Leaderboard, per-task heatmap, failure modes, and task taxonomy - [Paper (PDF)](https://hotinfra.org/2026/papers/hotinfra26-final71.pdf): Beyond Pass/Fail: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk - [HotInfra 2026](https://hotinfra.org/2026/): Workshop page (co-located with ISCA '26) - [GitHub repository](https://github.com/XuanmiaoG/InfraBench): Public task packages and project source surface - [Contributing guide](https://github.com/XuanmiaoG/InfraBench/blob/main/CONTRIBUTING.md): How to author and submit a task package - [Sitemap](https://infraben.ch/sitemap.xml): Crawlable URL list for the project site ## Optional - [Google Scholar (Yuan Gao)](https://scholar.google.com/citations?user=oD9j2NMAAAAJ&hl=en): Lead author publication profile - [Gemini Academic Program](https://ai.google.dev/gemini-api/docs/gemini-for-research): API credit sponsor for research evaluation - [CloudLab](https://www.cloudlab.us/): Compute testbed used for experiments - [CHTC](https://chtc.cs.wisc.edu/): Task interview partner - [UW–Madison DoIT](https://it.wisc.edu/about/division-of-information-technology/): Task interview partner - [ARA Wireless Living Lab](https://arawireless.org/): Task interview partner