Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
SMRTR summary
Terminal-Bench-Science is a Stanford benchmark testing AI agents on 70 real scientific tasks—spanning simulations, data analysis, and theorem proving—built with input from practicing scientists. The top model, Claude Opus 5, solved only 30% of tasks, highlighting major gaps in AI scientific capability.
SMRTR provides this summary for quick context. The original article belongs to Hacker News.
Read the original article