SEER: Enhancing Multimodal Action Grounding with Semantic UI Parsing
-
Updated
Nov 26, 2024 - Jupyter Notebook
SEER: Enhancing Multimodal Action Grounding with Semantic UI Parsing
CPU-friendly screen-to-JSON parser for AI desktop agents. Extracts pixel-perfect coordinates, OCR, UI semantics, cursor context, and structured screen state for reliable LLM-driven automation.
Add a description, image, and links to the screen-parsing topic page so that developers can more easily learn about it.
To associate your repository with the screen-parsing topic, visit your repo's landing page and select "manage topics."