Simple, Transparent Pricing
From generous free cloud access to unlimited self-hosting. Scale your AI infrastructure as you grow.
Flat-rate usage without token math
Start for free, scale up on the most affordable production tiers at $25/mo and $100/mo for production workloads.
- Start building completely free
- OpenAI-compatible endpoints
- Production analytics included
- Upgrade path to dedicated cloud hardware
Your Journey with Courier
Free Tier
Start on Courier Cloud at no cost to see how it works for you.
Pro Tier
Maxing out the free limits? Upgrade for higher usage and faster speeds.
Max Tier
For heavy cloud users who need the highest limits and fastest speeds.
Self-Managed
Ready for unlimited? Run Courier on your own Mac.
Fully Managed
Zero hassle. We manage your Mac hardware for you.
Courier Cloud
Free
The perfect way to start. Experiment with our Mac-optimized infrastructure at no cost.
- Generous rate limits
- Access to open-source models
- Models load on demand
Pro
Higher limits and faster speeds.
- Our fastest hardware
- Higher rate limits
- Lower latency and faster ttft speeds
Max
Maximum cloud capacity for power users and scaling applications.
- Everything in Pro
- Highest cloud rate limits
- Premium model access
- Priority cloud support
Self-Hosted License
Standalone
Install on your own M-series Mac. Fully featured and transferable. This is what Courier Cloud runs on.
- Unlimited tokens and requests
- Speed-optimized & production-ready
- Transferable between Macs
- Scout Pro included — Emissary + Scout Code
- Courier Relay — drive it from your iPhone
- Full privacy & data control
Managed
Send us your Mac, and we manage everything: backup power, internet, setup, and maintenance.
- All Standalone benefits
- No hardware upkeep hassle
- Backup power & redundant internet
- 24/7 technical maintenance
Production Ready Out of the Box
Our API platform goes far beyond model serving with dynamic memory management, industry-leading tool calling for agents, automatic hallucination detection, analytics, and the highest throughput on Apple Silicon
Flex + Static Deployment
Mix always-on and on-demand APIs on the same Mac.
Intelligent Memory Management
Never worry about OOMs again. LRU offload for flex models handles production spikes automatically.
Optimized Apple Silicon Speed
MLX-native stack with TurboQuant, KV caching, and robust batching gives you the best speeds and highest throughput on a single Mac.
Maximum Throughput
Robust batching for concurrency plus model pooling with a round-robin-styled API for redundancy. Courier maximizes throughput and speed on Apple hardware.
Automatic Hallucination Detection
Our system detects when models hang or hallucinate and automatically restarts them. You only lose one request instead of an entire queue, ensuring maximum uptime and reliability even with SLMs.
Industry-Leading Agent Capabilities
We built industry leading tool calling technology, making open source models on our platform the most performant for agentic workflows.
Real-Time Analytics
Track requests, tokens, latency, and model usage from day one.
Built-in Whisper API
Never pay for audio transcription again. Transcription and translation endpoints for automation pipelines.
Models Load on Demand
Limited memory? No problem. With the flex API you can load 10 models on one Mac Mini! Models load on-demand and intelligently negotiate memory with each other so you can use many different models in one constrained memory pool.
Always-On Production Endpoints
Static APIs stay loaded 24/7 for minimum latency. Pair static hot paths with flex overflow models for the best cost and performance mix.
Curated Cloud Models We Ship Today
Six production agents across lite, balanced, and frontier tiers — plus embeddings, image generation, and speech. Benchmarked, maintained, and ready on Courier Cloud.
Lite Agents
Best for classification, routing, extraction, and simple RAG agents where speed and cost matter most.
Gemma 4 E2B
Gemma 4 · omni
Our preferred compact Gemma 4 omni on Courier Cloud — faster than E4B with full multimodal support. Ideal for routing, fast sub-agents, and lightweight agentic workflows on memory-constrained devices.
Ternary Bonsai 4B
Bonsai · omni
The brains of a 4B with the speed of a 1B — a tiny, responsive lite agent for routing, extraction, and high-volume utility work.
Balanced Agents
The production default — Gemma 4 26B A4B for fast generalist work and Ternary Bonsai 27B for coding. Lower cost and fewer usage restrictions than frontier.
Gemma 4 26B A4B
Gemma 4 · omni
A fast, powerful generalist — chat, RAG, multimodal, and broad-knowledge tasks at balanced-agent cost with fewer usage restrictions than frontier.
Ternary Bonsai 27B
Bonsai · omni
Our balanced coding agent — strong on agentic coding and tool-using workflows at lower cost, though not quite as performant as frontier Qwen3.6.
Frontier Agents
Our two flagship agents — Gemma 4 31B for frontier generalist work and Qwen3.6 27B for specialist coding. Highest capability with tighter usage than balanced agents.
Gemma 4 31B
Gemma 4 · omni
Our frontier generalist agent — complex chat, multimodal reasoning, and broad-knowledge tasks at maximum quality.
Qwen3.6 27B
Qwen3.6 · omni
Our specialist frontier coding agent — built for the hardest coding agents and multi-step tool-using workflows.
Embeddings
Turn text into vectors for search, retrieval, clustering, and memory systems.
Qwen3 Embedding 8B
Qwen3 · text-embedding
A strong open-source embedding model — high retrieval quality for production RAG pipelines and semantic search.
Image Generation
Create images from text prompts for creative workflows, product visuals, and generated media.
FLUX.1 schnell
FLUX · text-image
Compact, fast image generator for product visuals, creative drafts, and generated media.
Speech
Transcription, translation, and audio understanding for automation pipelines.
Whisper
OpenAI · audio-text
Built-in speech-to-text for transcription and translation — available on every Courier Cloud plan.
Select multiple models for different tasks (e.g., coding, vision, and general chat). As your user base grows, you will see increased latency and degradation in user-experience if multiple models are not utilized.
Performance is determined by model quantization and available VRAM (Video Memory). Reasoning diminishes as quantization drops, possibly leading to hallucinations and other unintended side-effects.
- 4-bit: Maximum speed, lower VRAM
- 8-bit: Balanced speed and logic
- 16-bit: Maximum reasoning capability
Parameters are the internal variables the AI learns during training. A 30GB model has more "knowledge" than an 8GB model.
The context window is the amount of text (tokens) the AI can "remember" during a conversation or process in a single request.
Dynamic Memory Management
Courier offers 2 model serving options to maximize memory efficiency, Flex and Static
Flex models load into memory upon request and unload after 5 minutes of inactivity.
• Enables running multiple large models on limited hardware
• Dynamic memory allocation
• Only the largest flex model counts towards VRAM requirements
Static models stay loaded in memory at all times, providing instant response.
• Instant availability, no load time
• Continuous memory occupancy
• Each static model adds directly to total VRAM requirements
Start With a Use Case
Pre-configured flex stacks — memory is calculated from the largest flex model loaded at once.
What do you need AI for?
Select the primary functions for your self-hosted AI setup
Select Your Models
Choose the AI models to include in your platform (Filtered by your use cases)
No models selected. Add models to your platform to continue.
Hardware Recommendation
Need multi-device clustering or a custom setup? Book a free consultation