COLM 2026Computer-use agents
Real components.
Scalable evaluation.
A repeatable way to turn UI libraries into short, verifiable tasks. Test specific interactions, vary their context, and see how agents perform.
Paper benchmark · LLM-generated, human-verified
Move “Reports” to the top.
Drag a row. Or use Space to pick up, arrow keys to move, and Space to drop.
The task checks success automatically.
Built from the UI libraries behind real interfaces
From components to tasks
A repeatable way to build diagnostic tasks.
Start with a component, then vary its goal, layout, density, and surrounding context. AI-assisted construction produces task specifications, pages, and success checks; human validation checks usability and records reference trajectories.
- 01
Choose a component
Use real implementations from Ant Design, MUI, and Mantine.
- 02
Specify the task
Define the goal, scene, and target state in a structured specification.
- 03
Build and verify
Generate the interface and a programmatic success check.
- 04
Validate with people
Check that the task works and record human steps and time.
Explore the interactions
Short tasks. Specific skills.
Look beyond an overall score.
Inspect performance by component and observation/action interface, with human steps and time as a reference. Results describe the evaluated models and settings; new models can be measured on the same tasks.