mantismantis

A coding agent for any model.

The mantis CLI reads, edits and runs your code on the model you pick, local or hosted.

An SDK to build your own.

Tools, MCP, sessions and sub-agents in a few lines of Python. The same code runs on every model.

The mantis CLI adding a DELETE endpoint to a FastAPI app on qwen3-coder, with a syntax-highlighted diff

Runs on

OllamavLLMllama.cppTGITogetherFireworksGroqOpenRouterCerebrasClaudeOpenAIGeminiGrokModal
Deploy

Run any open model on your own GPUs.

Deploy GLM, Kimi, DeepSeek or Qwen to the GPU cloud you already use, and your agent is running on it in minutes.

localhost:8788/#deploy
The mantis dashboard Deploy page with a running open-model deployment
  1. 1
    Pick a model

    Search any open model. mantis checks it will fit before you spend anything.

  2. 2
    Pick a GPU

    See which GPUs fit, what they cost per hour, and choose one.

  3. 3
    Code with it

    When it boots, it is a model in your CLI and SDK. Stop it from the same page.

Deploys to
RunPodModalHugging FaceDeepInfraBasetenVast.aiFireworks
How deploys work →
Dashboard

One place to run all your agents.

mantis serve opens a local dashboard for everything on your machine.

Every session, every project

Open any conversation with its model, context fill and cost. Resume it in the terminal.

localhost:8788/#sessions
The mantis dashboard listing sessions across projects with model, context and cost
Open the dashboard guide →
Built in

What every model gets.

1agent = Agent(model="qwen3-coder:30b")route(model)wire ollama /api/chatname shape → backendno base_url, no adapterOllama · localroutedAnthropicOpenAIxAIGoogleTogether

Pick any model

Write the model name. mantis works out where it lives and how to talk to it.

tools.py@tooldef read_file(    path: str,    limit: int = 200) -> str:    """Read a file."""    ...tool_call{  "name": "read_file",  "arguments": {    "path": "app.py",    "limit": 200  }}nativepromptedconstrainedclaude-opus-5 · gpt-5.4 → API tool_calls

Tools that just work

Any Python function becomes a tool, even for models never trained to call one.

slacksse · 8 toolsgithubhttp · 9 toolsplaywrightstdio · 21 toolspostgresstdio · 4 toolsfilesystemstdio · 11 toolsmantis53 tools · 5 serversmcp__github__get_issue

Every MCP server

Connect filesystems, GitHub, databases and browsers, or publish your own tools.

qwen3-235b-a22b · fp8needs ~240 GBH100 80GB×4 · 320 GBfits · 80 GB headroom$2.49/h×4 $9.96A100 80GB×4 · 320 GBfits · 80 GB headroom$1.64/h×4 $6.56L40S 48GB×4 · 192 GBshort 48 GB$0.99/h×4 $3.96235Bprovisionpull weightsloadreadyready https://qwen3-235b.modal.run/v1

Your own GPUs

Pick an open model and a GPU. mantis deploys it and points your agent at it.

session a3f9 · maint1 – t20compact t1–t1452.4k → 6.1k tokreadeditcheckpointbashtestfixtestsqlitemigratetestfork b71cresume--resume a3f9/fork try-sqlite/rewind t16/compact

Sessions that keep going

Resume yesterday's work or fork to try another approach. Long chats compact on their own.

$0.00$0.25$0.50t7t1t12max_usd=0.50stop · $0.47tracetimecostturn 72.41s$0.031model gpt-5.41.62s$0.029tool bash640ms$0.002tool read_file90ms—

A hard spending limit

Cap a run in dollars. Every call is traced with its tokens and cost.

Models

The best models, open and closed.

Open weights. Use a hosted API, or deploy them to your own GPUs from mantis.

Modelmodel=Runs on
GLM-5.3ZhipuOpenRouter · z.ai · your GPU
Kimi K3MoonshotOpenRouter · your GPU
DeepSeek-V4 ProDeepSeekOpenRouter · your GPU
MiniMax M3MiniMaxOpenRouter · your GPU
Qwen3.8 27BAlibabaOpenRouter · Cerebras · your GPU
gpt-oss-120bOpenAIGroq · Together · Fireworks · Cerebras

Any other model works too. See all models and backends

Get started

Up and running in a minute.

01
Install

One package gives you the CLI, the SDK and the dashboard.

02
Choose a model

Paste a key, pick a local model, or deploy one to your GPUs.

03
Start

Open the agent in any project, or the dashboard in your browser.