6 quotes from 1 episode on Latent Space, each with a timestamped link to the source.
6 quotes1 episode
The short version
Ari Weinstein states AI agents operate software efficiently by interacting directly with source code and structural data. Generating JavaScript allows models to batch interface actions and complete tasks faster than average human users.
Most interesting insights
Ari Weinstein states agents now reliably introspect what works and debug their own errors when tasks fail.
“I think the biggest delta that I see is before they could like reliably start tasks but then they would run into problems and now they're really good at debugging. They're really good at trying again introspecting what is and isn't working.”
Ari Weinstein, Latent Space · October 2026 · Watch at 7:58 ↗
Attachment menus reveal the raw text and accessibility representations operating behind visual interfaces.
“If you click on the attachment and you click on this like little tiny button in the top right, you can see the raw text and you see the raw accessibility representation.”
Taking a standard screenshot hides hyperlink destinations and truncates long calendar event titles. Ari Weinstein notes an appshot provides full context by supplying the complete data behind visual elements.
“If you take a screenshot of a web page that has a link, the screenshot doesn't include where the link goes. It doesn't include, you know, maybe you take a screenshot of your calendar, the event titles are truncated, you know, but when you take an appshot, it gives like the language model like full context about everything.”
Agents write JavaScript to execute multiple actions
Models run batches of interface operations simultaneously. Ari Weinstein observes agents now write and execute code to perform several actions at once.
“Now computer use often writes code. So if you actually look at it in codecs and you expand the tool calls manually, you can see that it's not just doing one action at a time. It's actually writing JavaScript code that it executes that that the computer executes to perform sometimes many actions at once.”
Ari Weinstein, Latent Space · October 2026 · Watch at 8:19 ↗
Screen reader technology helps models use computers
The underlying systems built to make interfaces accessible to humans also feed structured semantic hierarchies to language models. Ari Weinstein emphasizes this technology makes computer use possible for AI.
“The same technology that was invented for humans who want to use a screen reader technology, that technology is really helpful for them to be able to use computers. It's also really helpful for an LLM to be able to use computers.”
Early Computer Use systems trapped agents in slow loops: take a screenshot, scroll down, take another screenshot, and guess the next pixel coordinate.
Modern agent architectures bypass pure vision by inspecting accessibility trees and DOM structures directly, letting the model read an entire page in one shot.
Raw screenshots hide critical structured data from language models, including hyperlink URLs and truncated calendar text.
OpenAI Codex uses App Shots, triggered by tapping both Command keys, to capture the underlying operating system accessibility tree alongside visual data.
How we attribute quotes. Every quote was matched against the episode transcript, so the words and the timestamp are real (we trim filler words like "um", nothing else). The name comes from our written summary of the episode, and we use it only when a separate check of the captions finds that person on the episode. YouTube gives us no voice-by-voice transcript, so open the timestamp to hear who is talking. See a wrong name? Tell us and we fix or remove it.