Skip to content

API Documentation ​

An OpenAI-compatible chat completion LLM API to easily integrate AI into your applications.

Quick Start ​

💡 For intensive usage (Mammouth Code, Cline), choose Starter + around $50 in API credits rather than Expert.

➡️ Get your API key and credits.

With the Mammouth API directly ​

Generates a chat completion response based on your prompt.

python
import requests
url = "https://api.mammouth.ai/v1/chat/completions"
headers = {
    "Authorization": "Bearer YOUR_API_KEY",
    "Content-Type": "application/json"
}
data = {
    "model": "gpt-4.1",
    "messages": [
        {
            "role": "user",
            "content": "Explain the basics of machine learning"
        }
    ]
}
response = requests.post(url, headers=headers, json=data)
print(response.json())
javascript
const fetch = require("node-fetch");

async function callMammouth() {
  const url = "https://api.mammouth.ai/v1/chat/completions";
  const headers = {
    Authorization: "Bearer YOUR_API_KEY",
    "Content-Type": "application/json",
  };

  const data = {
    model: "gpt-4.1",
    messages: [
      {
        role: "user",
        content: "Create an example JavaScript function",
      },
    ],
  };

  try {
    const response = await fetch(url, {
      method: "POST",
      headers: headers,
      body: JSON.stringify(data),
    });

    const result = await response.json();
    console.log(result.choices[0].message.content);
  } catch (error) {
    console.error("Error:", error);
  }
}

callMammouth();
bash
curl -X POST https://api.mammouth.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4.1",
    "messages": [
      {
        "role": "user",
        "content": "Hello, how are you doing?"
      }
    ]
  }'

➡️ Get your API key and credits.

With OpenAI Library ​

python
import openai

# Configure the client to use Mammouth.ai
openai.api_base = "https://api.mammouth.ai/v1"
openai.api_key = "YOUR_API_KEY"

response = openai.ChatCompletion.create(
    model="gpt-4.1",
    messages=[
        {"role": "user", "content": "What are the benefits of renewable energy?"}
    ]
)

print(response.choices[0].message.content)

Response Format ​

Successful Response ​

json
{
  "id": "chatcmpl-123",
  "object": "chat.completion",
  "created": 1677652288,
  "model": "gpt-4.1",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm doing very well, thank you for asking. How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 19,
    "total_tokens": 31
  }
}

Streaming Response ​

When stream: true is set, responses are returned as Server-Sent Events:

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1677652288,"model":"gpt-4.1","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-123","object":"chat.completion.chunk","created":1677652288,"model":"gpt-4.1","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}

data: [DONE]
json
{
  "id": "gen-1767710235-3VtWd1SuI9ilIspBmeWG",
  "created": 1767710235,
  "model": "google/gemini-2.5-flash-image",
  "object": "chat.completion",
  "choices": [
    {
      "finish_reason": "stop",
      "index": 0,
      "message": {
        "content": "Here's a beautiful sunset over mountains for you!",
        "role": "assistant",
        "images": [
          {
            "image_url": {
              "url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAABAAAAAQACAI..."
            }
          }
        ]
      }
    }
  ]
}

Models & Pricing ​

mammouth-recommended is a shortcut to whatever model Mammouth currently considers the best for its price.

Point your requests at it and you'll always get our current pick, without having to keep track of new releases yourself.

  • Current pick: glm-5.3-flash, with minimax-m3 as the fallback. This will change over time as new models come out.
  • How to use it: call it exactly like any other model. Set the model to mammouth-recommended, or use the shortcuts mammouth or recommended.
  • Pricing: you pay the same as the underlying model, with no markup, so check that model's row in the table below.

All models ​

Prices may vary and not be up to date in this table, please refer to the Model Explorer for the complete list of models !

ModelInput ($/M tokens)Output ($/M tokens)
claude-fable-5.11050
claude-haiku-4-515
claude-opus-5-5420
claude-sonnet-5210
deepseek-v4-pro1.743.48
deepseek-v4.1-flash0.220.66
gemini-3.1-flash-image-previewimage/
gemini-3.1-pro-preview212
gemini-3.7-flash1.57.5
gemini-3.8-flash0.753.75
glm-5.31.44.4
glm-5.3-flash0.150.5
gpt-5.42.515
gpt-5.4-mini0.754.5
gpt-5.4-nano0.21.25
gpt-5.5530
gpt-6-astra1050
gpt-6-luna0.10.5
gpt-6-sol210
grok-4.726
kimi-k2.60.733.49
kimi-k3315
llama-4-maverick0.150.6
minimax-m30.31.2
mistral-medium-3-51.57.5
mistral-small-3.2-24b-instruct0.10.3
qwen3.7-plus0.41.6
qwen3.8-27b0.42.55
qwen3.8-flash0.150.47
sonar-deep-research28
sonar-pro315

A note on pricing

The prices listed here are upper bounds.

What you actually pay is sometimes lower, since the price depends on provider availability.

You'll never be charged more than the price shown.

Embeddings ​

Generate vector embeddings for text to use in semantic search, clustering, and other NLP tasks.

Embedding Models & Pricing ​

ModelInput ($/M tokens)
text-embedding-3-large0.13
text-embedding-3-small0.02

Embedding Example ​

python
import requests

url = "https://api.mammouth.ai/v1/embeddings"
headers = {
    "Authorization": "Bearer YOUR_API_KEY",
    "Content-Type": "application/json"
}
data = {
    "model": "text-embedding-3-large",
    "input": "Hello, world!"
}
response = requests.post(url, headers=headers, json=data)
print(response.json())

Embedding Response ​

json
{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "index": 0,
      "embedding": [0.0023, -0.0091, 0.0152, ...]
    }
  ],
  "model": "text-embedding-3-large",
  "usage": {
    "prompt_tokens": 4,
    "total_tokens": 4
  }
}

📜 Usage and cost are logged in your settings.

💡 We added aliases aligned with the Mammouth app to facilitate your model selection: if you write mistral, it will use mistral-medium-3.1.

Error Codes ​

CodeDescription
400Bad Request - Missing or incorrect parameters
401Unauthorized - Invalid API key
429Too Many Requests - Rate limit exceeded
500Internal Server Error - Server-side issue
503Service Unavailable - Server temporarily unavailable

Tracking cost ​

If you want to know how much credits has been spent on a key, use this API endpoint:

bash
curl -X GET "https://api.mammouth.ai/key/info" -H "Authorization: Bearer $MAMMOUTH_API_KEY"

Parameters ​

Required Parameters ​

ParameterTypeDescription
messagesarrayList of messages in the conversation
modelstringModel identifier to use

Optional Parameters ​

ParameterTypeDefaultDescription
temperaturenumber0.7Controls creativity (0.0 to 2.0)
max_tokensinteger2048Maximum number of tokens to generate
top_pnumber1.0Controls response diversity
streambooleanfalseReal-time response streaming

Our API also supports provider-specific parameters and features, such as reasoning effort, tool calling, function calling, and thinking. Availability depends on the provider and model.

Optimization Tips ​

Message Structure ​

json
{
  "messages": [
    {
      "role": "system",
      "content": "You are an AI assistant specialized in programming."
    },
    {
      "role": "user",
      "content": "How to optimize a for loop in Python?"
    }
  ]
}

Role Types ​

  • system: Sets the behavior and context for the assistant
  • user: Represents messages from the user
  • assistant: Represents previous responses from the AI

Migration from OpenAI ​

If you're already using OpenAI's API, migrating to Mammouth.ai is simple:

  1. Change the base URL from https://api.openai.com/v1 to https://api.mammouth.ai/v1
  2. Update your API key
  3. Keep all other parameters the same

OpenAI Python Library ​

python
import openai

# Before
openai.api_base = "https://api.openai.com/v1"
openai.api_key = "sk-openai-key"

# After
openai.api_base = "https://api.mammouth.ai/v1"
openai.api_key = "your-mammouth-key"

n8n, VS Code, Cline, OpenClaw, Make, CLI, etc. ​

You can use the Mammouth API with tools like n8n, VS Code, Cline, Make and more.

Make sure you are using the correct URL. If unsure, try each of them.

Tutorials on how to use the Mammouth API in your favorite tools ​

For automations:

For IDEs:

For CLI (Claude Code equivalent):

Or via Opencode

Other

​