A coding agent for any model.
The mantis CLI reads, edits and runs your code on the model you pick, local or hosted.
An SDK to build your own.
Tools, MCP, sessions and sub-agents in a few lines of Python. The same code runs on every model.

import asyncio
from mantis_agent import (
query, MantisAgentOptions, tool, AssistantMessage,
)
@tool
async def get_weather(city: str) -> str:
"""Get the current weather for a city."""
return f"{city}: 67°F"
async def main():
async for msg in query(
prompt="What's the weather in SF?",
options=MantisAgentOptions(
model="qwen2.5:1.5b", # local Ollama
tools=[get_weather],
max_turns=5,
),
):
if isinstance(msg, AssistantMessage):
for block in msg.content:
if hasattr(block, "text"):
print(block.text)
asyncio.run(main())Runs on
Run any open model on your own GPUs.
Deploy GLM, Kimi, DeepSeek or Qwen to the GPU cloud you already use, and your agent is running on it in minutes.

- 1Pick a model
Search any open model. mantis checks it will fit before you spend anything.
- 2Pick a GPU
See which GPUs fit, what they cost per hour, and choose one.
- 3Code with it
When it boots, it is a model in your CLI and SDK. Stop it from the same page.
One place to run all your agents.
mantis serve opens a local dashboard for everything on your machine.
Every session, every project
Open any conversation with its model, context fill and cost. Resume it in the terminal.





What every model gets.
Pick any model
Write the model name. mantis works out where it lives and how to talk to it.
Tools that just work
Any Python function becomes a tool, even for models never trained to call one.
Every MCP server
Connect filesystems, GitHub, databases and browsers, or publish your own tools.
Your own GPUs
Pick an open model and a GPU. mantis deploys it and points your agent at it.
Sessions that keep going
Resume yesterday's work or fork to try another approach. Long chats compact on their own.
A hard spending limit
Cap a run in dollars. Every call is traced with its tokens and cost.
The best models, open and closed.
Open weights. Use a hosted API, or deploy them to your own GPUs from mantis.
| Model | model= | Runs on |
|---|---|---|
| GLM-5.3Zhipu | OpenRouter · z.ai · your GPU | |
| Kimi K3Moonshot | OpenRouter · your GPU | |
| DeepSeek-V4 ProDeepSeek | OpenRouter · your GPU | |
| MiniMax M3MiniMax | OpenRouter · your GPU | |
| Qwen3.8 27BAlibaba | OpenRouter · Cerebras · your GPU | |
| gpt-oss-120bOpenAI | Groq · Together · Fireworks · Cerebras |
Any other model works too. See all models and backends
Up and running in a minute.
One package gives you the CLI, the SDK and the dashboard.
Paste a key, pick a local model, or deploy one to your GPUs.
Open the agent in any project, or the dashboard in your browser.