Courier

Courier Platform

Courier Platform

Build without limits on the Courier Platform.

Cloud and local APIs for agents, products, and workflows. Build immediately on Courier Cloud, or air-gapped on your own Mac with no usage limits. Same platform either way.

DocsGet in touch

Build on Courier Cloud

Flat-rate APIs on Mac Studios we run. Curated library, analytics, no install. Start free with Basic Usage, or $500/mo for MAX usage supporting teams and products.

Self-host Courier Platform

Unlimited local inference on a Mac you own. Same API. Commercial is $300/mo flat. No usage limits.

Flex and static, one pool

Always-on hot paths stay resident. Overflow loads on demand and LRU-offloads when the turn is done. Use 5 models on a Mac that only fits one, swapping between them as needed.

Memory pool48 GB unified

Qwen 32B

Static — resident

Gemma 4

Flex — on demand

Waiting for a flex request

StaticStays hot for the path you always hit.

FlexLoads for a turn, then LRU offloads.

OpenAI-compatible

Drop-in endpoints. Same software on Cloud or a Mac you own. Automatic ngrok tunneling for easy API setup

Apple Silicon throughput

MLX-native stack, batching, KV cache — benchmark throughput on a Mac.

Industry leading tool calling

The same engine as Courier. Built for tool loops, not chat demos.

Analytics included

Requests, tokens, latency, model usage — from day one.

SDKs

Drop-in where you already work

OpenAI-compatible endpoints, n8n community nodes, and Python SDKs.

  • OpenAI-compatible
  • n8n
  • Python SDKs

Open-weight agents on Courier beat typical local stacks on tool-calling reliability. Read the Ollama benchmark.

Flat-rate. Not per token.

Cloud starts Free with Basic Usage, then $500/mo for Teams and Products. Self-hosted unlimited at $300/mo. Full tables on pricing.

Real-Time Analytics
Loading production data...
Find the Right Mac for Your Use Case
Start from a preset or build your own flex stack, then see which Apple Silicon Mac fits your models.
Model Count

Select multiple models for different tasks (e.g., coding, vision, and general chat). As your user base grows, you will see increased latency and degradation in user-experience if multiple models are not utilized.

1-3 Models:Focused Setup
4-10 Models:Versatile Setup
10+ Models:Full Ecosystem
Throughput - Quantization & VRAM

Performance is determined by model quantization and available VRAM (Video Memory). Reasoning diminishes as quantization drops, possibly leading to hallucinations and other unintended side-effects.

  • 4-bit: Maximum speed, lower VRAM
  • 8-bit: Balanced speed and logic
  • 16-bit: Maximum reasoning capability
Model Size - Parameters

Parameters are the internal variables the AI learns during training. A 30GB model has more "knowledge" than an 8GB model.

Lite (1GB - 14GB):Fast, Efficient
Balanced (15GB - 50GB):Versatile, Strong
Frontier (50GB+):Advanced Reasoning
Context Window - Memory

The context window is the amount of text (tokens) the AI can "remember" during a conversation or process in a single request.

32k tokens:~50 pages of text
128k tokens:Full book length
1M+ tokens:Entire codebases

Dynamic Memory Management

Courier offers 2 model serving options to maximize memory efficiency, Flex and Static

Flex Models

Flex models load into memory upon request and unload after 5 minutes of inactivity.

• Enables running multiple large models on limited hardware

• Dynamic memory allocation

• Only the largest flex model counts towards VRAM requirements

Static Models

Static models stay loaded in memory at all times, providing instant response.

• Instant availability, no load time

• Continuous memory occupancy

• Each static model adds directly to total VRAM requirements

Feeling overwhelmed or unsure what to choose?

Let us help you figure it out.

Start With a Use Case

Pre-configured flex stacks — memory is calculated from the largest flex model loaded at once.

Courier
Chat agent, lightweight sub-agent, and embeddings for Pathfinder
All models use flex APIs
Coding Agent
Planner + implementer stack for agentic coding workflows
All models use flex APIs
Production Server
General-purpose production API with Gemma 4
All models use flex APIs

What do you need AI for?

Select the primary functions for your self-hosted AI setup

Agent
Tool-calling agents, chat, and multi-modal workflows
Image Generation
Generate images from text prompts
Embeddings/RAG
Semantic search, retrieval, and memory

Select Your Models

Choose the AI models to include in your platform (Filtered by your use cases)

No models selected. Add models to your platform to continue.

Hardware Recommendation

Infrastructure Requirements
Based on your model selection
Total VRAM Required7 GB
Recommended HardwareMac mini (M4, 16GB) — from $599

Need multi-device clustering or a custom setup? Book a free consultation