Courier Platform
Build without limits on the Courier Platform.
Cloud and local APIs for agents, products, and workflows. Build immediately on Courier Cloud, or air-gapped on your own Mac with no usage limits. Same platform either way.
Build on Courier Cloud
Flat-rate APIs on Mac Studios we run. Curated library, analytics, no install. Start free with Basic Usage, or $500/mo for MAX usage supporting teams and products.
Self-host Courier Platform
Unlimited local inference on a Mac you own. Same API. Commercial is $300/mo flat. No usage limits.
Flex and static, one pool
Always-on hot paths stay resident. Overflow loads on demand and LRU-offloads when the turn is done. Use 5 models on a Mac that only fits one, swapping between them as needed.
Qwen 32B
Static — resident
Gemma 4
Flex — on demand
Waiting for a flex request
StaticStays hot for the path you always hit.
FlexLoads for a turn, then LRU offloads.
OpenAI-compatible
Drop-in endpoints. Same software on Cloud or a Mac you own. Automatic ngrok tunneling for easy API setup
Apple Silicon throughput
MLX-native stack, batching, KV cache — benchmark throughput on a Mac.
Industry leading tool calling
The same engine as Courier. Built for tool loops, not chat demos.
Analytics included
Requests, tokens, latency, model usage — from day one.
SDKs
Drop-in where you already work
OpenAI-compatible endpoints, n8n community nodes, and Python SDKs.
- OpenAI-compatible
- n8n
- Python SDKs
Open-weight agents on Courier beat typical local stacks on tool-calling reliability. Read the Ollama benchmark.
Flat-rate. Not per token.
Cloud starts Free with Basic Usage, then $500/mo for Teams and Products. Self-hosted unlimited at $300/mo. Full tables on pricing.
Select multiple models for different tasks (e.g., coding, vision, and general chat). As your user base grows, you will see increased latency and degradation in user-experience if multiple models are not utilized.
Performance is determined by model quantization and available VRAM (Video Memory). Reasoning diminishes as quantization drops, possibly leading to hallucinations and other unintended side-effects.
- 4-bit: Maximum speed, lower VRAM
- 8-bit: Balanced speed and logic
- 16-bit: Maximum reasoning capability
Parameters are the internal variables the AI learns during training. A 30GB model has more "knowledge" than an 8GB model.
The context window is the amount of text (tokens) the AI can "remember" during a conversation or process in a single request.
Dynamic Memory Management
Courier offers 2 model serving options to maximize memory efficiency, Flex and Static
Flex models load into memory upon request and unload after 5 minutes of inactivity.
• Enables running multiple large models on limited hardware
• Dynamic memory allocation
• Only the largest flex model counts towards VRAM requirements
Static models stay loaded in memory at all times, providing instant response.
• Instant availability, no load time
• Continuous memory occupancy
• Each static model adds directly to total VRAM requirements
Start With a Use Case
Pre-configured flex stacks — memory is calculated from the largest flex model loaded at once.
What do you need AI for?
Select the primary functions for your self-hosted AI setup
Select Your Models
Choose the AI models to include in your platform (Filtered by your use cases)
No models selected. Add models to your platform to continue.
Hardware Recommendation
Need multi-device clustering or a custom setup? Book a free consultation