COLM 2026Computer-use agents

Real components.
Scalable evaluation.

A repeatable way to turn UI libraries into short, verifiable tasks. Test specific interactions, vary their context, and see how agents perform.

2,910 tasks97 component types14 families

Paper benchmark · LLM-generated, human-verified

Try the benchmark
Paper benchmark · v1
Ant DesignReordering

Move “Reports” to the top.

Drag a row. Or use Space to pick up, arrow keys to move, and Space to drop.

Loading the task…

The task checks success automatically.

Built from the UI libraries behind real interfaces

Ant DesignMUIMantine

From components to tasks

A repeatable way to build diagnostic tasks.

Start with a component, then vary its goal, layout, density, and surrounding context. AI-assisted construction produces task specifications, pages, and success checks; human validation checks usability and records reference trajectories.

  1. 01

    Choose a component

    Use real implementations from Ant Design, MUI, and Mantine.

  2. 02

    Specify the task

    Define the goal, scene, and target state in a structured specification.

  3. 03

    Build and verify

    Generate the interface and a programmatic success check.

  4. 04

    Validate with people

    Check that the task works and record human steps and time.

Look beyond an overall score.

Inspect performance by component and observation/action interface, with human steps and time as a reference. Results describe the evaluated models and settings; new models can be measured on the same tasks.