Semantic, not pixels
Agents read labelled fields and typed rows — not screenshots they have to squint at. No vision model in the loop, no coordinate guessing.
A multi-window desktop app, readable and operable over MCP — while your users keep clicking buttons, in sync, on the same projection.