LessWrong AI
2026-09-28 08:04 UTC
By Ana Leonescu
USR-0152-20260928-community-fo-4da2af03
Why do models *really* fail on HLE tasks?
I recently attended Generality Labs ' Inspect Evals Data Viz Hackathon, and spent the day using Inspect AI and its offspring, Scout, with a simple goal in mind - generate a new plot of a new or existing benchmark. Many thanks to the organisers and to my team mates, Jeff Mohl and Valerie Griffiths (the Overfit and Overcaffeinated team), for a fantastic time, learning some new tricks on using the Inspect suite. Here I'm presenting the two (!) plots we got in the span of ~ 6 hours (more like 4 hours as it took us a while to agree on what we actually want to spend the day on - arguably a harder task than its execution). [... 2 hours later... ] Our initial idea was to run a few models on a subset of Humanity's Last Exam (HLE) , and undertake an extensive failure mode analysis to understand the current gaps and where different capabilities x harnesses fail or succeed. We did this using Inspect Scout - a framework for an LLM-as-judge that analyses the models' outputs and assigns a dominant feature that lead to the answer failing or winning. Then, we rerun a subset of tasks and analysed how changing the harness affects the prevalence of failures - given the time constraints, we only modified the harness to include access to web search. All code is available here. Initial failure mode analysis We started with 200 randomly sampled HLE tasks, relatively balanced across domains, and ran three GPT models on the same fixed subset. This gave us 600 model attempts in total, of which 525* we…
I recently attended Generality Labs ' Inspect Evals Data Viz Hackathon, and spent the day using Inspect AI and its offspring, Scout, with a simple goal in mind - generate a new plot of a new or existing benchmark. Many thanks to the organisers and to my team mates, Jeff Mohl and Valerie Griffiths (the Overfit and Overcaffeinated team), for a fantastic time, learning some new tricks on using the Inspect suite. Here I'm presenting the two (!) plots we got in the span of ~ 6 hours (more like 4 hours as it took us a while to agree on what we actually want to spend the day on - arguably a harder task than its execution). [... 2 hours later... ] Our initial idea was to run a few models on a subset of Humanity's Last Exam (HLE) , and undertake an extensive failure mode analysis to understand the current gaps and where different capabilities x harnesses fail or succeed. We did this using Inspect Scout - a framework for an LLM-as-judge that analyses the models' outputs and assigns a dominant feature that lead to the answer failing or winning. Then, we rerun a subset of tasks and analysed how changing the harness affects the prevalence of failures - given the time constraints, we only modified the harness to include access to web search. All code is available here. Initial failure mode analysis We started with 200 randomly sampled HLE tasks, relatively balanced across domains, and ran three GPT models on the same fixed subset. This gave us 600 model attempts in total, of which 525* we…
Full article content could not be extracted automatically. Read the original below.