BrowserVM by browserscale
Your code, beside the page.
Not inside the page, where the site can see it. Not at the far end of a socket, paying a round trip for every question it asks. BrowserVM runs your script in an isolate of its own inside the browser, with a direct line into the exact document you are touching — so a cross-origin iframe is read with a dot, an element can be handed straight to a click, and no part of it names a frame.
// A cross-origin <iframe> in a process of its own. Inside the page,
// reading contentDocument here returns null — the same-origin policy
// forbids it, and that DOM is not even in this process.
const checkout = window.document.getElementById('checkout').contentDocument;
// A live element out of that other document, handed straight to a
// command. The whole walk can be the argument.
await browser.click(checkout.getElementById('accept-terms'));
await browser.fill(checkout.querySelector('input[name=code]'), '4711');
// Real property access on the real objects, in the other renderer.
// This is an assignment, not a message to a copy.
const code = checkout.querySelector('input[name=code]').value;
// Selectors are still here when you want them, and they refine.
await browser.click(css('#submit').visible().steady(500));
return code;
per step into the document, so a loop over a hundred table rows is milliseconds rather than a conversation.
script injected, compiled or evaluated in the page. A site that rewrites its own internals sees a read and a call the same way: not at all.
of frame, same-origin or cross-origin, in its own process or not, reached by reading a property — and it costs the same.
Where the script runs
There have only ever been two places to put automation. Inside the page, where it is fast but visible to everything the page does — a global here, a rewritten prototype there, and the site knows. Or outside the browser, where it is invisible to the page but pays a network round trip for every question it asks, which is why remote automation is written as few, coarse commands rather than as code.
BrowserVM is a third place. Your script runs in an isolate of its own, in its own sandboxed process, with a direct line into the exact document you are touching. It is not in the page, so the page cannot see it. It is not outside the browser, so it does not pay for distance. It is beside the page.
Fine-grained automation stops being expensive, and unobservable automation stops being coarse.
One consequence is worth stating early, because everything else on this page follows from it: work that was only ever practical as a single big blob of injected JavaScript can now be written as ordinary code, step by step, with real control flow — and the page still sees nothing.
It is not a reduced mode beside the real API, either. Every command the session has is callable from in here, on the same session and in the same run — the catalogue further down goes through them.
Every boundary is a property
Every other tool makes you name things. A frame handle, a target session, something that has to be re-resolved after a navigation. It is the single biggest source of brittleness in real automation: the checkout field is in an iframe, the iframe is on another origin, so it lives in another process, so you first go and find it, then address everything through it, and you keep that bookkeeping correct for the rest of the script.
Here you walk into a frame the way you walk into anything else — by reading a property. It works on same-origin children, on cross-origin ones, on children in their own process, and at any depth, because the element you read from may itself have come out of a child document.
const f = window.document.getElementById('cross');
// The child's own document, from another origin in another process.
f.contentDocument.getElementById('i').textContent;
// And really the child's realm — not the parent pretending.
f.contentWindow.origin;
// Its globals are writable, and the write lands over there.
f.contentWindow.sessionFlag = 'ok';
And because an element is a target, the hop and the action are one expression. Nothing in it names a frame: not a locator, not an option, not a helper.
await browser.click(
window.document.getElementById('cross')
.contentDocument.getElementById('cb'));
Depth does not stack up, which is the first thing anyone driving an iframe tree asks. Reaching a document is a one-off; finding a frame inside it does not care how deep it sits. Measured, a grandchild two out-of-process frames down costs about 1.4× a step in the main frame — flat, not per level.
Live objects, not snapshots
A value that comes back is the real object, not a serialized copy of it. Property reads, property writes and method calls all happen on that object in its own frame, so there is no bridge in the middle that could answer from a stale picture.
const input = window.document.querySelector('#x');
input.value = 'typed directly'; // a real assignment
const box = input.getBoundingClientRect(); // runs where the element lives
const d = window.document;
const li = d.createElement('li');
li.textContent = 'added from the outside';
d.querySelector('ul').appendChild(li); // both objects are remote
Identity means what it says. d.body === d.body is true, and so is d.getElementById('x') === d.getElementById('x'): one object read twice comes back as one object, whichever route it was reached by. Deduping a collected list, or asking whether an element is the focused one, gives the right answer rather than one that merely looks right.
Objects from two different frames are never mixed silently — handing a call an object from elsewhere is refused outright. Frames are crossed by walking into them, not by passing their objects around.
And when a page navigates out from under a script that is holding references to it, those references throw at the line that touched them, as ordinary exceptions. They cannot silently resolve against whatever loaded next, which means a walk across a moving page is something you can wrap in a try and carry on from.
Nothing is injected
Reading or writing a property adds nothing to the page. No injected global, no altered prototype, no evaluated source, nothing for the page to hook. A page-defined accessor still runs, because it is part of the object — that is the page's own code doing the page's own work, exactly as it would if a person had touched the field.
Method calls are not interposable either. A call names its function and the engine invokes it directly. Nothing on that path is expressible in page JavaScript, so a page that instruments its own internals to catch automation sees a method call the same way it sees a read: not at all. The test for this is a page that deliberately booby-traps the two functions such a call would historically have gone through, and then gets driven anyway, with the traps never firing once.
Invisibility here is a property of where the code sits, not a set of tricks that a new detection probe can undo.
Your console.log goes to you, never to the page's console. And when you do want to run something in the page, that is a separate, explicit call — browser.evaluate(...) — visible by construction, because you asked for it.
Arm, act, await
The script keeps executing while a command is pending, so the arm-then-trigger pattern collapses into a single script with no race between two remote calls. This is the shape a wire protocol structurally cannot offer, because between arming and triggering it has a network.
// Armed before the thing that causes it, awaited after.
const pending = browser.waitForBlobImage({ timeoutMs: 6000 });
await browser.click(css('#generate-captcha'));
const shot = await pending;
// The same shape for anything on the wire.
const token = browser.waitForResponse('*/api/challenge*');
await browser.click(css('#verify'));
const body = (await token).response.body;
That is how you reach a captcha image that only ever exists as a blob URL, and a token that only ever exists as a response. For a challenge drawn into a canvas, browser.readCanvas reads the pixels even when the canvas is tainted, without ever calling into the page to do it. And __wrc.shadow(host) reaches into a closed shadow root natively, so the parts of a page that are meant to be sealed off are addressable too.
What a step costs
One step into the document costs around thirty microseconds, where the same step from outside the browser is a network round trip. So roughly six reads per table row means a hundred rows in about fifteen milliseconds, and a loop that would be unthinkable as a remote conversation is merely unremarkable here.
The honest other half: a walk is not a query. Thirty thousand reads is still under a second, but the same thirty thousand reads written as one browser.evaluate that loops inside the page take about a millisecond. So walk when you are crossing boundaries, acting, or holding on to identity — and reach for evaluate the moment the job is bulk text extraction. Both are one line away from each other, in the same script.
One path, start to finish
A whole flow in the order a real run uses it: the interruptions handled declaratively, one wait covering every branch the page has, a payment field inside the processor's own cross-origin iframe, a token taken off the wire, a hundred rows read as ordinary code, and the click that matters guarded rather than hoped at.
// 1. The consent bar is not a step in the flow, so it is not in it.
await browser.addReaction(css('#accept-all'));
await browser.navigate('https://shop.example/checkout',
{ timeoutMs: 30000 });
// 2. Four ways this page can go. One deadline, one branch.
const step = await browser.waitAny([
css('#payment-form'),
css('#login-email'),
css('#queue-position').visible(false).steady(0),
css('[data-error]'),
], { timeoutMs: 25000 });
if (step.index === 3) return { ok: false, reason: 'rejected at entry' };
if (step.index === 1) {
await browser.fill(css('#login-email'), account.email);
await browser.fill(css('#password'), account.password);
// Armed before the click that causes it, awaited after — no race,
// because the script never left the browser in between.
const auth = browser.waitForResponse('*/api/session*');
await browser.click(css('#sign-in'));
console.log('sign-in returned', (await auth).response.status);
await browser.wait(css('#payment-form'), { timeoutMs: 20000 });
}
// 3. The card fields live in the processor's own cross-origin iframe,
// in its own process. Read into it with a dot; this script never
// mentions a frame id.
const pay = window.document
.querySelector('iframe[name=pay]').contentDocument;
await browser.fill(pay.querySelector('[name=number]'), card.number);
await browser.fill(pay.querySelector('[name=cvc]'), card.cvc);
// 4. A hundred summary rows, read as ordinary code. Each step is
// microseconds, so the loop is not a conversation.
const rows = window.document.querySelectorAll('#summary tr');
const lines = [];
for (let i = 0; i < rows.length; i++) {
const c = rows[i].children;
lines.push({ sku: c[0].textContent, qty: +c[1].textContent });
}
// 5. The click that matters. It scrolls, settles, checks the pixel,
// aims around anything over it — and if it still cannot land it says
// what stopped it instead of clicking something else.
try {
await browser.click(css('#place-order').visible().steady(750));
} catch (e) {
console.error(e.message); // names the intercepting element
return { ok: false, reason: e.message, lines };
}
await browser.wait(css('#order-number'), { timeoutMs: 40000 });
// 6. Take the signed-in identity along, so the next run starts here.
return {
ok: true,
lines,
order: window.document.getElementById('order-number').textContent,
session: await browser.getAuthSession(),
};
Three of those steps are not BrowserVM features at all — the standing reaction, the four-way wait and the click that refuses come from the engine underneath, and they work exactly the same when you drive the session from your own process. That machinery is catalogued further down, because it is what makes a script written in here this short.
Why an agent writes this better
The hardest part of machine-written automation is not the clicking. It is bookkeeping: which frame was that element in, which handle is still valid, which session am I addressing, what do I re-resolve after this navigation. A model that gets any of that subtly wrong produces a script that works once and then fails in a way nobody can read.
Here there is no bookkeeping to get wrong. An element is a target and answers for its own frame. A walk is an expression, so the model writes what it means in one line. Failures land at the line that caused them, with the engine's own message — and where a click or a wait fails, that message names the element responsible, which is exactly what a model needs in order to repair its own script rather than start over.
And the loop an agent actually runs — look at the page, decide, act, look again — is one round trip per cycle instead of dozens, because the deciding happens next to the looking. getObservation gives the model a page it can reason about; the same script then acts on it without leaving the browser.
The engine it sits on
Inheriting the whole session API matters more than it sounds, because these commands are not thin wrappers around DOM calls. They are the parts of the browser that were rebuilt — and they are the reason the flow in section 07 is as short as it is.
Three of them carry most of the weight, so they are worth spelling out before the catalogue.
Waiting is reported, not polled
Each document that a condition concerns watches for it itself and pushes a wake-up the moment it starts holding, so a match arrives within a frame or two rather than on the next tick of a timer — and a frame in its own process reports for itself. A condition means visible, reachable and still for half a second by default, where reachable is decided by the same hit test a click would do, so an element behind a modal is not reported as ready. Several conditions in one call give you the branch the page actually took. And when nothing matches in time you get a per-condition diagnosis: not_found, found_hidden, found_occluded with the element that covered it, or pending_steady for something that was there but still moving.
A click that lands, or explains itself
One call locates the element, scrolls it into view through nested scrollers and up the frame chain, waits until its bounds stop moving, moves the pointer there along a human path — real acceleration and braking rather than a jump, jitter that settles as it arrives, and a landing point that is not the exact centre — and then verifies the exact pixel the way a real mouse event is routed, which is what makes it correct across process boundaries. If something covers the point it re-aims at the largest exposed patch of the element; failing that it steps clear of the whole overlay once, which collapses hover menus, and re-checks. If it still cannot land it refuses rather than clicking the wrong thing, and returns the intercepting element: tag, id, class, text, bounds, computed pointer-events, stacking order, whether it is pinned, and whether it swallows clicks while invisible.
A script in here is short because the hard decisions were already made one layer down.
Interruptions are handled without you
A standing reaction is an instruction the browser carries out on its own: when this appears anywhere in the page, click it — or click something beside it, like the close button. It fires only while the pointer is idle, so it slots into the gaps of an action that is already retrying; a click blocked by a modal therefore lands, because the reaction dismissed the modal while the click was still working its way in. It is one-shot, watches every frame including frames created later, and survives navigation.
The catalogue
Targets are interchangeable throughout: a live element, a css() selector, a js() expression, a node handle from an earlier result, or bare coordinates.
- browser.wait · waitAny
- One condition or several, first match wins, with the frame and node that matched — plus the diagnosis above when nothing does.
- browser.click · moveTo · drag · scrollTo
- The full path described above, plus pointer movement on its own, dragging with that same machinery on the source, and scrolling through nested containers and frame chains. Options for button, double clicks, press-only and release-only, and the settle window.
- browser.fill · type · insertText · key
- Filling focuses with the click core, optionally clears, puts the caret at the end and types on the keyboard layout of the session's region with human cadence — stopping and naming the thief if something steals focus mid-value. Typing is deliberately loose for one-time-code boxes; inserting text is one commit with no key events; keys can be pressed and released separately.
- browser.addReaction · listReactions · removeReaction
- Standing instructions, what is still armed, and how to take them back down.
- browser.selectByIndex · ByValue · ByText
- Picks an option in a select, with or without firing events.
- browser.waitForRequest · waitForResponse · modifyRequest
- Arm before the action, await after it: method, URL, headers, status and body. Mark a pattern to be aborted and it is answered with an empty 200 instead. Headers and bodies can be rewritten in flight.
- browser.setBlockList · setStaticPaths · loadHTML
- Drop requests across the whole context, serve repeated static assets from a cache instead of the network, or answer a whole navigation with your own response without touching the network.
- browser.getAuthSession · setAuthSession
- A signed-in persona as a single object — account, refresh token and the device-bound sessions that keep it alive. Import it into a fresh context and it comes up signed in.
- browser.cookies · storage · setProxy
- The whole cookie jar and local storage per origin, read and written as data, and the context's exit changed at runtime.
- browser.getObservation
- The page reduced to what a model can act on: one line per element with role, type, name, live value, label and flags — disabled, required, off-screen, clickable — frame by frame, bounded by a token budget. Invisible-but-still-interactive elements are kept and marked, because those are the ones that trip a script up.
- browser.getDOM · screenshot · highlightNode
- The tree as structured data, a frame of the page rendered on the session's own GPU, and a border drawn on an overlay above the page — no DOM change, no layout effect, invisible to the page.
- browser.readCanvas · waitForBlobImage · inspectAtPosition
- Canvas pixels past the origin-clean restriction, the bytes of the next image the page mints as a blob URL, and what is actually on top at a point.
- browser.evaluate · browser.command
- Run something in the page on purpose, or call any engine command directly — including ones without a named wrapper yet. Nothing is walled off from a script.
All of it is equally available from your own process through the Go and TypeScript SDKs, the CLI and the MCP server. The platform page covers that side, including the parts BrowserVM does not change: real-GPU fingerprints, managed proxies, captcha solving, live WebRTC control and thousands of isolated sessions in parallel.
Runs you can watch, and end
A run either hands back its value when it is done, or answers immediately with an id and leaves the outcome to an event. The second is what long work wants: the script keeps going without the caller, which is how automation outlives the process that started it.
Everything a script prints arrives while it is still running, each line tagged with its run and stamped when it was written rather than when it was received. A one-shot run also carries its whole log back in the reply, so a single call does not need a subscription just to read its own output. You can list what is still in flight — useful when you arrive to find work you did not start — and you can end a run by id or all of them at once. That reaches both cases that matter: a script parked on an await, and a script spinning in a tight loop with no await at all.
$ browserscale run <session-id> checkout.js
$ browserscale run <session-id> -e "return window.document.title"
$ cat probe.js | browserscale run <session-id> -
out, err := b.RunScript(ctx, source) // wait for the value
run, err := b.StartScript(ctx, source, onEv) // leave it running
runs, err := b.ListScriptRuns(ctx) // what is still going
What it does not do
A page that oversells is a page you stop trusting, so: the isolate is bare. There is no host environment around your script — no timers, no fetch, no DOM constructors of its own. To wait, you wait on something real in the page, which is what the waits are for.
Remote objects are objects, not collections: a remote node list is indexed, not iterated, so spread and for…of do not apply to it. And every property access is a real step, so the guidance in section 06 is a rule and not a suggestion.
None of that is in the way of the work. It is the shape of a tool that lives where it lives.
Get early access
BrowserVM runs inside browserscale sessions, alongside real GPU-backed rendering, managed proxies, portable identities and live remote control. Nothing to install, and nothing to run on your side.
It is being opened account by account while we work through it with the first users. Tell us what you want to automate, and we switch it on for your account — everything else on the platform is open to every account today.