For LLMs
The complete site text, in one place. Projects, experience, contact, and both full local blog posts.
Abdul Rashid
Founding engineer and engineering lead at Reclaim Protocol.
Work spans product and infrastructure, from the experience people use to the systems that keep it running.
Outside work, I explore spatial computing, especially tangible user interfaces that connect physical objects with digital systems. Current interests include AR, projection mapping, and computing beyond the screen.
Selected projects
Popcorn
An open-source browser cloud built for scale, security, and performance.
Browser fleets scale from 10 to 10,000 pods, with a live view designed for phones. Touch, typing, gestures, and reconnection make remote sessions feel like a local browser.
Built the regional fleet and mobile viewer together, keeping browsers warm for fast allocation and adding capacity as demand grows. Sessions combine end-to-end encryption, confidential hardware, and hardware attestation for a verifiable browser environment.
30-day snapshot
- Browser sessions
- 529,739
- Browser hours
- 30,603
- Median allocation
- 643 ms
Playfield
An iPhone, a projector, and a wall that plays back.
A spatial computing experiment that turns a wall into games controlled by your hands. Defend a city, cook imaginary soup, or draw in the air. The phone tracks your hands and runs the game; the projector brings it into the room.

All projects → — AI-agent research, Sentilytics, and Tap Tap ML.
Blog
Why I’m building Playfield
From EyeToy with friends to games on a wall. A story about the joy of playing with your body.
Can you leave your agent alone for ten hours?
Building evaluation harnesses that let AI agents work with more autonomy.
All posts → — Browser clouds, faster mobile proofs, and notes from building.
Experience
Reclaim Protocol (YC W21)
–PresentFull-time
- Founding Engineer
- –Present · Remote
- Engineering Lead
- –Present
- Software Engineer
- –
Work spans the full stack: native and web SDKs, developer tools, verification agents, and the infrastructure behind them. Led projects from implementation through evaluation and production operations.
- Led the Reclaim app’s migration from React Native to Flutter, alongside work on mobile SDKs, web SDKs, and developer tools.
- Built Popcorn’s browser cloud and mobile live viewer, bringing together regional fleet operations, elastic scaling, end-to-end encryption, and hardware attestation.
- Built and evaluated AI agents for creating verification flows. An autoresearch harness compares models and approaches, guides experiments, and fine-tunes smaller models, reducing agent cost by 97% and making execution 30% faster.
- Led the gnark migration, moving mobile proof generation from a WebView to native Go. Average proof time fell from about 40 seconds to 4–5 seconds.
- Built Glassbreak, a Slack-based system for managing cloud permissions and tracking access for audits.
- Built internal tools for engineering and operations: an AI review agent, a log viewer, and triage and Snitch bots for investigating production issues.
Full experience → — SaaS Labs, Amazon, Procedure, and earlier roles.
Let’s talk
Interested in something I’m building, have an idea to share, or just want to chat about something nerdy? Say hello.
abdulrreshamwala@gmail.comBlog
Notes from building infrastructure, giving AI agents more autonomy, and exploring spatial computing.
Why I’m building Playfield
From EyeToy with friends to games on a wall. The story behind Playfield, inspired by Party Fowl, Folk Computer, and the joy of playing with your body.
Can you leave your agent alone for ten hours?
How evaluation harnesses let AI agents work with more autonomy, with case studies from XPath-Go, Popcorn, and our provider-creation agent.
Why We Built Popcorn: An Attestable Browser Cloud for Reclaim
Building our own browser cloud: confidential computing, a regional fleet, and a browser experience that feels natural on a phone.
Turbocharged Zero-Knowledge Proofs for Mobile
Moving mobile proof generation from a WebView to native Go and gnark, cutting average proof time from about 40 seconds to 4–5 seconds.
Experience
Engineering across product and infrastructure, from payments and event-driven systems to browser clouds.
Career
Reclaim Protocol (YC W21)
–PresentFull-time
- Founding Engineer
- –Present · Remote
- Engineering Lead
- –Present
- Software Engineer
- –
Work spans the full stack: native and web SDKs, developer tools, verification agents, and the infrastructure behind them. Led projects from implementation through evaluation and production operations.
- Led the Reclaim app’s migration from React Native to Flutter, alongside work on mobile SDKs, web SDKs, and developer tools.
- Built Popcorn’s browser cloud and mobile live viewer, bringing together regional fleet operations, elastic scaling, end-to-end encryption, and hardware attestation.
- Built and evaluated AI agents for creating verification flows. An autoresearch harness compares models and approaches, guides experiments, and fine-tunes smaller models, reducing agent cost by 97% and making execution 30% faster.
- Led the gnark migration, moving mobile proof generation from a WebView to native Go. Average proof time fell from about 40 seconds to 4–5 seconds.
- Built Glassbreak, a Slack-based system for managing cloud permissions and tracking access for audits.
- Built internal tools for engineering and operations: an AI review agent, a log viewer, and triage and Snitch bots for investigating production issues.
SaaS Labs
–Software Development Engineer 2 · Full-time
Modernized legacy backend systems, migrating PHP monoliths to JavaScript microservices and redesigning callback processing for lower latency and greater scale.
- Moved callback processing from the legacy cron-based system to an event-driven architecture. Callbacks could be processed as events arrived, removing the wait for a scheduled run.
- Reduced callback latency by more than 20× and made processing horizontally scalable, with capacity to handle 100× the legacy system’s traffic.
- Improved coding and refactoring guidelines across teams to make the evolving services easier to maintain.
Amazon
–Software Developer · Full-time
- Developed merchant registration for Amazon India, Amazon Pay, and Fresh, and a peer-to-merchant payment flow for Amazon Pay.
- Built responsive interfaces and a reusable component library for the frontend.
- Handled design, implementation, unit testing, and deployment through AWS Pipelines.
- Managed cloud infrastructure, automated deployments, and monitored production systems with CloudWatch.
Procedure
–SDE II · Full-time · Mumbai
Built backend infrastructure and real-time monitoring for last-mile delivery, with a focus on high throughput and low latency.
- Designed and deployed the backend infrastructure supporting delivery operations and analytics.
- Built real-time monitoring and analytics to follow delivery activity as it happened.
- Developed load-testing tools for both traffic bursts and sustained load, helping evaluate throughput, latency, and system behavior under pressure.
Genius Consulting
–Software Engineer · Full-time
Designed, developed, and deployed an automated stock-trading system. Built scrapers to collect and warehouse live data from multiple platforms, and worked on stock-performance evaluation and recommendations.
Metanoia Technologies
–Software Engineer · Full-time
Built proofs of concept for mobile and web projects to explore product ideas and evaluate technology choices. Planned the software development lifecycle and set up and maintained cloud deployment infrastructure.
Dimensionless Tech.
–Data Scientist · Internship
- Created hands-on learning modules, including signature classification.
- Worked on Baggage AI, a computer vision platform for detecting objects in baggage X-ray images.
- Extended the internal certificate-generation system, refactored older code into reusable components, and added website payment processing.
Education
Rizvi College of Engineering
2017–2021B.E. in Computer Engineering
Projects
Browser infrastructure, AI agents, and experiments in computing beyond the screen.
Popcorn
An open-source browser cloud built for scale, security, and performance.
Browser fleets scale from 10 to 10,000 pods, with a live view designed for phones. Touch, typing, gestures, and reconnection make remote sessions feel like a local browser.
Built the regional fleet and mobile viewer together, keeping browsers warm for fast allocation and adding capacity as demand grows. Sessions combine end-to-end encryption, confidential hardware, and hardware attestation for a verifiable browser environment.
30-day snapshot
- Browser sessions
- 529,739
- Browser hours
- 30,603
- Median allocation
- 643 ms
Playfield
An iPhone, a projector, and a wall that plays back.
A spatial computing experiment that turns a wall into games controlled by your hands. Defend a city, cook imaginary soup, or draw in the air. The phone tracks your hands and runs the game; the projector brings it into the room.

Autoresearch harness for AI agents
The goal was to make AI agents faster and cheaper to run, while reducing the need to test every change by hand.
An autoresearch harness automates experiments, compares models and approaches, and uses measured results to fine-tune smaller models. Repeatable checks measure task completion, cost, and speed, giving the next experiment a clear direction.
- Lower agent cost
- 97%
- Faster execution
- 30%
Sentilytics
An AI-powered API for sentiment and emotion analysis in video.
Built with college friends for a competition, Sentilytics combines video frames, audio, and transcripts to understand a clip from multiple angles. It captions scenes, detects facial emotions, tracks people across the timeline, and analyzes sentiment in speech and text.
A Flask API and React interface bring the results together, using OpenCV and FFmpeg for video processing, TensorFlow models, and IBM Watson for speech-to-text.
Tap Tap ML
A no-code path from datasets to deployed models.
Create datasets, train custom machine learning models, and deploy inference to serverless functions.
More on GitHub — Browse the repositories and see what else I’m building.Can you leave your agent alone for ten hours?
By Abdul Rashid · 6 min read

Getting a PR from a reported issue is already something AI models can handle pretty well. But give them a long horizon task and they start to struggle, lose direction, and take shortcuts. If we have to keep looking over the agent’s shoulder to check its work, we are still the bottleneck.
If we want to give agents more autonomy, we need to create a harness. By harness, we mean an evaluation loop around the task, beyond tools like Codex and Claude Code. It should emulate how we would verify the work and guide the agent: try it, find what is wrong, and give it enough feedback to know what to do next. That lets it keep moving without waiting for us after every change.
Some case studies from what we have been working on.
XPath-Go
Moving to the TEE stack meant we needed an XPath library in Go. The agent got parts of it working and the Go tests were passing, but the results differed from our JS version. A simple comparison test gave us a way forward: give both versions the same input and compare what comes back. This grew into a suite comparing our Go library against jsdom, checking nodes, text, namespaces, and source locations.
That made a big difference. The agent could see exactly where the two implementations disagreed and keep working through those differences. I could give it a goal, go off for around ten hours, and come back to a much better implementation. The recorded run passed 888 out of 888 compatibility cases. There are still things outside that test coverage, but we now had a concrete way to measure progress without checking each fix ourselves. Results
Popcorn
The goal with Popcorn was to build the best mobile webview. That meant testing how it performs across a wide variety of devices, operating systems, and websites. Even figuring out what to test is hard: every website has different interactions, and those can behave differently depending on the screen size, OS, or keyboard. Turning all of that into a fixed set of automated tests is a lot of work. So we built a flow where the agent explores interactions, creates test cases and repeatable test pages, then compares the same interaction directly in the device browser and through Popcorn.
We wanted the testing to be close to how we would do it ourselves. So it uses native taps and gestures at coordinates and captures the actual screen. If an input is hidden behind the keyboard, we need to see that. Knowing that the browser thinks the input is focused does not tell us whether someone can use it. We test across iOS, Android, and multiple keyboards.
We also started putting parts of the checking into code. Pixel diffs let us repeat visual checks with defined thresholds, without spending AI tokens looking at every screenshot on every run. The agent can spend its time finding new cases and investigating failures. Harness
An AI agent improving an AI agent (AIception?)
For our provider-creation agent, the goal was to have AI improve it, try new providers, and compare different models and approaches. The harness was an evaluation setup with 22 education portals under our control. With known accounts, test data, and expected results, we could run each variation through the same tasks and see what worked.
Creating a provider was only part of the test. It had to return the logged-in user's complete name, not just their first name, and the proof could not leak their username or password. The generated injection also had to detect login correctly, keep the verification overlay hidden before login, and show it afterwards. These were the things we would check ourselves, now built into the evaluation.
The agent still found ways to bend the rules. In one experiment, it added a provider fixer during evaluation. The task was to create a provider and check whether it could replay successfully. Fixing it along the way made the result look better, but did not tell us whether the original provider worked. So the repair step was removed: the harness pins the exact version created and replays it in a fresh session with AI disabled. Both creation and replay have to pass. Playback contract
More autonomy still needed some human nudges. Sometimes the agent would keep trying variations of the same approach when it needed to look somewhere else. Sharing what we already knew about the system and giving it access to sister repos helped it do that. Some failures led to Portal filtering valid JSON responses; others involved the TEE getting an empty response body. With that broader context, it could investigate the actual problem across repos instead of trying to fix everything inside the agent repo.
More trust, more autonomy
Agents will keep getting better, but if we are still checking everything they do because we don’t trust them, we will miss out on a lot of those gains. Our time is better spent improving the checks, sharing context, and giving a new direction when needed. Letting the harness handle repeated testing builds the trust to give agents more autonomy. It also lets us run more agents in parallel without every result queuing up for us to check.
Why I’m building Playfield
A wall, an iPhone, and the joy of playing with your body.
By Abdul Rashid · 5 min read
Some of my favourite childhood memories are of playing EyeToy with friends. A camera plugged into the PlayStation, a game on the TV, and suddenly our bodies were the controllers.
There’s a particular feeling you get from that kind of play. You reach for something and the game responds. You move a little too much, look a little ridiculous, and your friends laugh. The fun spills out of the screen and into the room. Even watching someone else play becomes part of it.
EyeToy made that feel possible with a little camera. What stayed with me was how much fun it was to move around and play together.
That feeling came back
Years later, playing Party Fowl brought it back. It uses a device’s camera to turn body movements into controls for wonderfully silly games. There was that familiar feeling again: moving, reacting, and laughing at what the game was asking us to do.
It reminded me how much I like games that give you a reason to get up. The movement is part of the pleasure. So is being in the same room as the people you’re playing with.
That’s the feeling behind Playfield, an experiment I’m building with an iPhone, a projector, and games you play with your hands. The game goes on the wall. Your hands do the rest.
Let the room join in
Spatial computing has interested me for a long time, especially mixed reality, augmented reality, and tangible interfaces. I like the possibility of reaching into a digital experience through the space and objects around us.
Folk Computer was a big inspiration. It uses cameras and projectors to bring computation onto physical surfaces and objects. Paper can become an interface. Moving something on a table can change what a program does. The computer gets a place in the room, alongside the people using it.
What drew me in was how much room that leaves for experimentation. An interface can be something you touch, rearrange, or share with another person. Seeing Folk made me want to explore that direction myself.
Playfield starts with a small piece of it: a wall that responds to your hands. Something you can walk up to, understand by trying, and play together.
A smaller setup
Interactive walls and floors have existed for years. Systems like LUMOplay already turn projected surfaces into places to play. Their recommended setups combine a computer, a projector, and a dedicated 3D camera.
That made me curious about what a phone could take over. A modern iPhone brings a camera, the compute to run hand-tracking models, and, on supported models, LiDAR depth sensing into one device. AI vision models can locate hands in the camera image; depth can help with interactions that need to know how close a hand is to a surface.
For this experiment, the phone handles tracking and runs the games. The projector makes them big enough to play on a wall. LiDAR is used in the wall-touch mode; the other games use camera-based hand tracking.
There’s still a projector to connect and a phone to mount. But being able to try this without a separate computer and tracking camera makes it much easier to keep experimenting.
A wall with a few new rules
The phone first looks for markers projected at the corners of the play area. That lets it map what the camera sees to where things appear on the wall. Once calibration is ready, the markers disappear and the game takes over.
In Meteor Keeper, your hands steer paddles to bounce meteors away from a city. In Counter Crew, you grab tomatoes, chop them with your hand, cook imaginary soup, and serve it. There’s a drawing mode for leaving trails across the projection, and a physics playground for grabbing and dropping a ball.
The games are deliberately small. They’re ways to find out which interactions feel good: reaching, grabbing, opening your hand to let go. When the hands leave view, the games can pause. Even an imaginary kitchen should let you take a break.
What I want to build toward
Playfield is still an experiment. A lot of the work is making the tracking, calibration, and gestures feel natural enough that you can pay attention to playing. A hand movement should do what you expect, without making you think about the camera watching it.
The goal is a setup that’s easy to bring out, with games that invite people to join in. There’s plenty to explore beyond these first few games, but this feels like a good place to start.
EyeToy gave me some of my favourite memories with friends. Party Fowl reminded me how good that kind of play can feel. Folk opened up more possibilities for what the space around us could do.
Now there’s a little city on my wall, and it needs saving.