sypherin

@altronis/tokenmark-mcp

Community sypherin
Updated

TokenMark MCP server and CLI: measured local-LLM speeds and model picks for Strix Halo, DGX Spark and Mac

@altronis/tokenmark-mcp

An MCP server that gives Claude (and other agents) real local-LLM benchmark data + hardware-aware model recommendations from TokenMark.

Ask "what should I run on a Strix Halo for coding?" and the agent answers from measured configs (decode tok/s, quant, backend), each with a source link. It never invents numbers.

Add to Claude Code

claude mcp add tokenmark -- npx -y @altronis/tokenmark-mcp

Add to Cline

Cline CLI:

cline mcp add tokenmark --yes -- npx -y @altronis/tokenmark-mcp

Cline in VS Code: add the tokenmark entry from the JSON below to cline_mcp_settings.json. Step-by-step notes for agents are in llms-install.md.

Add to any MCP client

Run the server over stdio:

npx -y @altronis/tokenmark-mcp

Or in a client config:

{
  "mcpServers": {
    "tokenmark": {
      "command": "npx",
      "args": ["-y", "@altronis/tokenmark-mcp"]
    }
  }
}

Tools

  • tokenmark_recommend: { hardware, tasks?, prefer?, limit? } → ranked model + best-config picks (5 by default, 10 max) with measured tok/s + why.
  • tokenmark_configs: { model?, hardware?, limit? } → tracked benchmark configs, fastest first (25 by default, 50 max; matched gives the full count), with the run mode (speculative: true for MTP/DFlash/draft-model runs).
  • tokenmark_hardware: { platform? } → the hardware catalogue (Strix Halo, Gorgon Halo, DGX Spark, Mac Max/Ultra). With a platform: chip specs with a source per value, the boxes that ship it, Singapore prices.
  • tokenmark_search: { term } → matching models/configs.
  • tokenmark_submit: { repo, source?, note? } → queues a GitHub/Hugging Face repo with benchmark numbers for human review.
  • tokenmark_submission_status: { id } → where a submission is.

A speed someone measured themselves, or hardware missing from the catalogue, goes through the signed-in form at https://tokenmark.app/submit.

Data is pulled live from https://tokenmark.app (override with TOKENMARK_URL). Zero runtime dependencies.

The same data in your terminal

The CLI lives in cli/ of this repo:

npx @altronis/tokenmark-cli recommend --hardware "Strix Halo" --tasks coding

The files in bin/ are the published builds, made from the TokenMark tracker that runs https://tokenmark.app.

MIT licensed. Data aggregated from public community benchmarks with attribution.

MCP Server · Populars

MCP Server · New

    unibaseio

    membase-ai

    Membase SDK: drop-in memory infrastructure for AI agents and apps — Python client, CLI, MCP server and TypeScript client, hosted or on your machine.

    Community unibaseio
    fingentic

    Revolut MCP

    This is the Revolut MCP server

    Community fingentic
    nekyialabs

    Resonant Mind

    Persistent cognitive infrastructure for AI systems. 28 MCP tools — semantic memory, emotional processing, identity continuity, and a subconscious daemon. Built on Cloudflare Workers.

    Community nekyialabs
    kirill-markin

    Nibomo

    AI-powered flashcards app built for serious daily study on iOS, Android, and the web. Use it to prepare for exams, learn vocabulary, memorize technical terms and facts, improve your material with AI, and review with spaced repetition.

    Community kirill-markin
    audiojs

    audio

    High-level audio manipulations

    Community audiojs