This page is also a Node script. A model with a shell and a KVM snapshot needs no browser. Give it the capture at its real size, and say the size of your own screenshot of that screen — only you know it:
node cadena-lince-r3.html shot.png --frame 1456x819 the manifest, in the frame you click in: your own screenshot's size
node cadena-lince-r3.html shot.png --frame mine.png the same, read off a picture: your own screenshot
node cadena-lince-r3.html shot.png no frame said: one is assumed, and the answer says ASSUMED
node cadena-lince-r3.html shot.png --frame capture in the capture's own pixels, as r1 answered
node cadena-lince-r3.html shot.png --out frame.png and the frame as a picture: what you see is what you click
node cadena-lince-r3.html shot.png --zoom 100,55,200,115 --out zoom.png
look closer at that box: x,y stay full-frame
node cadena-lince-r3.html shot.png --zoom 37 the same, around zone 37
node cadena-lince-r3.html shot.png --find "Compose" best match first line as: x y
node cadena-lince-r3.html shot.png --at 148,82 what stands at that point of the frame
node cadena-lince-r3.html shot.png --space kvmd each line also as the KVM takes it: kvmd:X,Y
node cadena-lince-r3.html shot.png --space capture each line also in real pixels: px:X,Y
node cadena-lince-r3.html --frame 1456x819 --at 148,82 --space kvmd
no image: one frame point into pointer units
node cadena-lince-r3.html --help all options
The frame is declared, never guessed in silence. A harness decides what size of picture its model is handed, and two harnesses give two sizes for the same screen. So the size comes from the model's side: --frame WxH, a picture of that size, or LINCE_FRAME=WxH in the environment, said once for a whole session. Where nobody says, the usual frame under a 1568 px picture limit stands in and the second line of the answer reads ASSUMED, with the sum that puts it right.
Zoom to identify, click in full-frame. A zoom is cut from the capture's own pixels and drawn as large as a model can be shown; its rulers read full-frame x and y, and each zone carries its manifest number. The model looks at the zoom to see which icon is which, then clicks with the frame coordinates it already had. It never counts in real pixels and never in the zoom's own.
In a page, the engine is window.LINCE: await LINCE.fromBlob(blob) returns the zones, LINCE.manifest(result, {frame:[1456,819], space:"kvmd"}) the text, LINCE.find(result, "Compose", {frame:[1456,819]}) the point, LINCE.at(result, [148,82], {frame:[1456,819]}) what stands there, LINCE.picture(result, {frame:[1456,819], zoom:[100,55,200,115]}) the pixels to look at. The page takes the same words on its address: #frame=1456x819&space=kvmd&zoom=100,55,200,115. document.documentElement.dataset.state goes idle → busy → done, dataset.frame holds the frame in use and dataset.frameSaid whether it was declared or only assumed, and the manifest sits in #out.
Reading is deterministic: the same pixels give the same zones, and the same zone numbers. Small print under about 7 px cap height, print laid over photographs, and print cut off by a window edge are left alone rather than guessed — which is why the capture should be the real screen and not the model's small frame of it.
Node runs the file as plain CommonJS. Inside a project whose package.json says "type": "module" it refuses the .html name: copy the file to lince.cjs and run that.