Interesting new benchmark called RLE for AI agents in robotics engineering. Looks like GPT-6 Astra has the top mark for now.
# | Model / Harness | RLE Index |
1 | GPT-6 Astra OpenAI · Codex | 73.4 |
2 | Claude Opus 5 Anthropic · Claude Code | 53.4 |
3 | GPT-5.6 Sol OpenAI · Codex | 48.7 |
4 | DeepSeek-V4.1-Flash DeepSeek AI · Claude Code | 38.9 |
5 | Gemini 3.7 Flash Google · Antigravity CLI | 26.8 |
Can Your AI Engineer a Robot? [9/17, Harvard University News]
The new benchmark contains 48 tasks that test whether general-purpose coding agents can perform the kinds of engineering work required to build and operate robotic systems. Tasks range from developing control and perception algorithms to designing robot hardware and interfaces. RLE-Bench is open source and publicly available, allowing researchers to evaluate new AI systems on a common set of tasks.
Jim
--
You received this message because you are subscribed to the Google Groups "RSSC-List" group.
To unsubscribe from this group and stop receiving emails from it, send an email to rssc-list+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/rssc-list/b9dfbe4c-bb94-4257-bb9b-3f3d3db039e3n%40googlegroups.com.