Kvxw1105

lean-computer-use-mcp

Community Kvxw1105
Updated

Low-context, state-safe MCP facade over Open Computer Use for inexpensive agent models

lean-computer-use-mcp

Low-context, state-safe MCP facade over Open Computer Use for inexpensive agent models such as GPT-5.6 Luna.

Status: M1 verified against the real Windows upstream (cu_find_app, cu_observe, metrics, cu_act stale-rejection and real-action paths, including a JianYing subtitle resize). V2 vision fallback and vision=auto LLM escalation are live. Record & Replay (demonstrate a workflow once, replay it cheaply) is implemented as CLI commands. Not yet recommended for production use.

Why this project exists

Open Computer Use works, but every snapshot includes a screenshot and every action returns a full refreshed UI state. On Windows we measured:

Payload Size
Default get_app_state tree text ~54,000 characters
Compact READ tree text ~2,300 characters
Screenshot (Base64) ~405,000 characters, unchanged between presets

A skill can reduce how often a model observes, but it cannot remove screenshots, action-returned full states, or duplicated tool schemas from the model's context. This project puts a bounded proxy between the model and the upstream server so the model sees only what it needs to complete the task.

Measured on the real desktop (ChatGPT window, 2026-08-05): the default upstreamsnapshot costs ~437,779 model-visible characters (55,543 text + 382,236 imageBase64) and 460 nodes; the facade's cu_observe returns an 820-characterpayload with 3 controls and no image, a 99.8% reduction in model-visiblecontext. See docs/BENCHMARKS.md for the full table andreproduction commands.

Procedural memory (atomic components)

Beyond whole-task replay, compile --library and recall learn atomiccomponents (e.g. jianying::click::button::font-size) and task templates,then compose new tasks from old building blocks. Replay feeds results back:successes raise popularity and teach effects, failures raise staleness.refine lets the model curate the library (aliases, merges, descriptions,template generalizations) with a human-reviewed apply step.See docs/MEMORY.md.

Record & Replay

Demonstrate a workflow once, then replay it with far less context:

lean-computer-use record --app JianYing --out recordings/font-size.json
lean-computer-use compile --in recordings/font-size.json --out-dir skills/recorded/subtitle-font-size
lean-computer-use replay --in recordings/font-size.json --run

The recorder captures mouse/keyboard events plus periodic element snapshots(no screenshots), compiles an editable, intent-based SKILL.md (like theofficial macOS-only Codex Record & Replay), and replay re-locates targets inthe live tree - coordinates are only a fallback for custom-rendered UIs.See docs/RECORDING.md.

Architecture

flowchart LR
    Model[Low-cost model e.g. Luna] --> Skill[lean-computer-use-luna skill]
    Skill --> Facade[lean-computer-use-mcp]
    Facade --> Cache[Local state + image cache]
    Facade --> Upstream[open-computer-use MCP/CLI]
    Upstream --> Windows[Windows UIA / screenshot]

The facade owns:

  • compact, query-relevant accessibility output instead of full trees;
  • state_id-based freshness and stale-state rejection;
  • local screenshot caching and on-demand cropping;
  • delta summaries after actions instead of full refreshed states;
  • per-call metrics for honest before/after cost measurement.

Repository layout

docs/            DESIGN, PROTOCOL, SECURITY, BENCHMARKS
src/             Python MCP server (incl. record/compile/replay CLI)
tests/           unit tests and fixtures
skills/          Codex skill that drives the facade
benchmarks/      benchmark scenario definitions
config/          example agent configuration

Development

git clone https://github.com/<you>/lean-computer-use-mcp.git
cd lean-computer-use-mcp
uv sync --all-extras
uv run pytest

Run a demo server with a fake upstream client (no desktop access):

uv run lean-computer-use serve --fake

Documentation

  • Design
  • Protocol
  • Security
  • Benchmarks

License

MIT

MCP Server ยท Populars

MCP Server ยท New

    drakulavich

    Kesha Voice Kit

    Give your tools a voice โ€” speech to text and back, 25 languages, up to ~19ร— faster than Whisper. On your machine.

    Community drakulavich
    lobu-ai

    Lobu โ€” Open-source backend for AI teammates

    Open-source control plane and runtime for organisational agents: shared company context, isolated execution, approvals and MCP.

    Community lobu-ai
    minipuft

    Claude Prompts MCP Server

    Wolfflow: Model Context Protocol (MCP) server for reusable prompt templates, multi-step workflow chains, and quality gates. Compose agentic workflows with an operator syntax; export as native skills to Claude Code, Cursor, OpenCode, and Gemini CLI.

    Community minipuft
    docmancer

    Docmancer

    Find out what your coding agents already know. Docmancer indexes the memory, rules, and instructions Claude Code, Codex, Cursor, and Gemini wrote on your machine, then carries the durable parts to every agent. Local-first, MIT.

    Community docmancer
    lineai-intelligence

    codelogic-mcp-server

    An MCP Server to utilize Codelogic's rich software dependency data in your AI programming assistant.