Key Takeaways
- Raw screenshots hide critical structured data from language models, including hyperlink URLs and truncated calendar text.
- OpenAI Codex uses App Shots, triggered by tapping both Command keys, to capture the underlying operating system accessibility tree alongside visual data.
- Accessibility screen-reader APIs supply structured semantic hierarchies that let models understand interactive states without bloating token budgets.
- Models perform best on the exact data representations they were trained on, making custom scraping pipelines prone to distribution errors.
Most teams building computer-using agents make the same mistake. They take full-resolution screenshots, pipe the raw image into a multimodal vision model, and expect the model to click the right button.
It fails constantly. When you take a standard screenshot of a web page, the pixels show blue underlined text, but they do not show the destination URL. When you screenshot a crowded calendar, the event titles get clipped by the visual box. The model has to guess what sits behind the pixels.
Ari Weinstein and the team at OpenAI took a different approach with Codex and ChatGPT. Instead of treating the computer screen purely as an image, they built App Shots. By extracting the operating system's native accessibility tree and DOM metadata, App Shots give models clean, structured context while protecting token budgets.
Why Raw Pixels Blind Your Agents
Multimodal vision models are expensive and lossy when applied to user interfaces. A visual screenshot consumes thousands of image tokens while dropping the exact metadata an agent needs to execute reliable actions.
Weinstein pointed out the obvious failure modes of pure image capture: “If you take a screenshot of a web page that has a link, the screenshot doesn't include where the link goes. It doesn't include, you know, maybe you take a screenshot of your calendar, the event titles are truncated, you know, but when you take an appshot, it gives like the language model like full context about everything.”
In Codex, users trigger an App Shot by pressing both Command keys at the same time. This captures the active application state instantly. As Weinstein explained, “If you click on the attachment and you click on this like little tiny button in the top right, you can see the raw text and you see the raw accessibility representation.”
By passing both visual structure and text trees, the agent avoids hallucinating hidden labels or clicking ambiguous coordinates.
Turning Screen Readers into Model Interfaces
The engineering breakthrough behind App Shots did not require inventing a new operating system protocol. It relied on accessibility infrastructure built decades ago for assistive technology.
Screen readers for visually impaired users already parse UI hierarchies, track focused elements, expose accessibility labels, and resolve interactive trees. That exact structural stream turns out to be ideal for language models.
The hard part is compression. Raw accessibility trees from macOS or Windows are massive dumps filled with irrelevant layout coordinates and empty nodes. OpenAI spent substantial engineering time filtering this tree down: “We've put a lot of work into dumping it out, but also making it token efficient. Doing it efficiently, there's like a bit of an art to it.”
The Distribution Trap of Custom Wrappers
Founders building custom agent frameworks often spend weeks writing their own browser DOM scrapers or desktop accessibility parsers. That work frequently backfires.
Frontier models are fine-tuned on specific representations of UI trees. If your custom parser formats accessibility nodes differently from how OpenAI or Anthropic formatted them during training, the model operates out of distribution. It stumbles on simple multi-step workflows.
Unless you are pretraining your own vision-action models, your agent stack should stay as close to the model provider's native format as possible.
What to Do With This
Inspect the exact UI payload your agents send to your model provider this week. If you are sending raw full-screen images, replace that pipeline with an accessibility tree dumper or a DOM snapshot that strips invisible nodes, inline styles, and SVG paths. Verify that URLs and truncated text fields are exposed as plain text attributes before the payload enters your prompt.