# Can you leave your agent alone for ten hours?

By Abdul Rashid · 6 min read

![A hand-drawn rider guiding a horse with a harness](<https://abdulrashid.dev/blog/assets/horse-harness-doodle.png>)

Getting a PR from a reported issue is already something AI models can handle pretty well. But give them a long horizon task and they start to struggle, lose direction, and take shortcuts. If we have to keep looking over the agent’s shoulder to check its work, we are still the bottleneck.

If we want to give agents more autonomy, we need to create a harness. By harness, we mean an evaluation loop around the task, beyond tools like Codex and Claude Code. It should emulate how we would verify the work and guide the agent: try it, find what is wrong, and give it enough feedback to know what to do next. That lets it keep moving without waiting for us after every change.

Some case studies from what we have been working on.

## XPath-Go

Moving to the TEE stack meant we needed an XPath library in Go. The agent got parts of it working and the Go tests were passing, but the results differed from our JS version. A simple comparison test gave us a way forward: give both versions the same input and compare what comes back. This grew into a suite comparing our Go library against jsdom, checking nodes, text, namespaces, and source locations.

That made a big difference. The agent could see exactly where the two implementations disagreed and keep working through those differences. I could give it a goal, go off for around ten hours, and come back to a much better implementation. The recorded run passed **888 out of 888 compatibility cases**. There are still things outside that test coverage, but we now had a concrete way to measure progress without checking each fix ourselves. [Results](https://github.com/reclaimprotocol/xpath-go/blob/main/docs/COMPATIBILITY.md)

## Popcorn

The goal with Popcorn was to build the best mobile webview. That meant testing how it performs across a wide variety of devices, operating systems, and websites. Even figuring out what to test is hard: every website has different interactions, and those can behave differently depending on the screen size, OS, or keyboard. Turning all of that into a fixed set of automated tests is a lot of work. So we built a flow where the agent explores interactions, creates test cases and repeatable test pages, then compares the same interaction directly in the device browser and through Popcorn.

We wanted the testing to be close to how we would do it ourselves. So it uses native taps and gestures at coordinates and captures the actual screen. If an input is hidden behind the keyboard, we need to see that. Knowing that the browser thinks the input is focused does not tell us whether someone can use it. We test across iOS, Android, and multiple keyboards.

We also started putting parts of the checking into code. Pixel diffs let us repeat visual checks with defined thresholds, without spending AI tokens looking at every screenshot on every run. The agent can spend its time finding new cases and investigating failures. [Harness](https://github.com/reclaimprotocol/popcorn-oss/blob/main/images/minimal-vnc-desktop/mobile-harness/README.md)

## An AI agent improving an AI agent (AIception?)

For our provider-creation agent, the goal was to have AI improve it, try new providers, and compare different models and approaches. The harness was an evaluation setup with 22 education portals under our control. With known accounts, test data, and expected results, we could run each variation through the same tasks and see what worked.

Creating a provider was only part of the test. It had to return the logged-in user's complete name, not just their first name, and the proof could not leak their username or password. The generated injection also had to detect login correctly, keep the verification overlay hidden before login, and show it afterwards. These were the things we would check ourselves, now built into the evaluation.

The agent still found ways to bend the rules. In one experiment, it added a provider fixer during evaluation. The task was to create a provider and check whether it could replay successfully. Fixing it along the way made the result look better, but did not tell us whether the original provider worked. So the repair step was removed: the harness pins the exact version created and replays it in a fresh session with AI disabled. Both creation and replay have to pass. [Playback contract](https://github.com/reclaimprotocol/education-portal-evals-infra/blob/main/docs/provider-playback-validation.md)

More autonomy still needed some human nudges. Sometimes the agent would keep trying variations of the same approach when it needed to look somewhere else. Sharing what we already knew about the system and giving it access to sister repos helped it do that. Some failures led to Portal filtering valid JSON responses; others involved the TEE getting an empty response body. With that broader context, it could investigate the actual problem across repos instead of trying to fix everything inside the agent repo.

## More trust, more autonomy

Agents will keep getting better, but if we are still checking everything they do because we don’t trust them, we will miss out on a lot of those gains. Our time is better spent improving the checks, sharing context, and giving a new direction when needed. Letting the harness handle repeated testing builds the trust to give agents more autonomy. It also lets us run more agents in parallel without every result queuing up for us to check.
