Carraggon CarraWorks
CarraWorks is Carraggon's inference and fine-tuning platform: serverless endpoints, dedicated GPU deployments, managed training and embeddings — one API, from first prototype to enterprise scale.
Time to first token
Seconds, no cold starts
Serving
Serverless or dedicated GPUs
Billing
Per token, or per GPU second
Deployment
Cloud, VPC or on-prem
Building with open models shouldn't mean building an inference team. CarraWorks handles the serving, the tuning and the scaling — you keep the weights, the data and the choice of model.
The platform
Start serverless, move to dedicated capacity when volume justifies it, tune your own models, run your own RL loops, and route spend down with Nexus.
01
Serverless Inference
Get started in seconds with per-token pricing, zero setup and no cold starts.
Pay per token with high rate limits and postpaid billing. Standard, Priority and Fast tiers let you trade cost against latency per workload — same API, one line to switch.
02
On-Demand Deployments
Pay per GPU second for faster speeds, higher rate limits and lower cost at scale.
Dedicated capacity billed per second. Pin a model to your own GPUs for predictable latency, private networking and unlimited throughput once volume justifies it.
03
Managed Training
Customize open models with your own data with minimal setup.
Supervised fine-tuning (SFT), preference tuning (DPO) and reinforcement fine-tuning, fully managed. Upload a dataset, pick a base model, and deploy the result as a LoRA add-on in minutes.
04
Embeddings & Rerankers
Retrieval quality is the ceiling on agent quality.
Hosted embedding and reranking endpoints priced per million input tokens, sized from compact 150M models to 8B-class encoders for the hardest enterprise corpora.
05
RL Rollouts
The only dedicated rollout inference for teams running their own RL.
Rollout-optimized inference built for reinforcement learning loops: massive parallel sampling, in-flight weight sync and per-second GPU billing so your trainer never waits on the sampler.
06
Nexus
Increase token usage, reduce cost. An open-weight model layer behind your harness.
Nexus sits behind your existing agent harness and routes each call to the cheapest open-weight model that clears your quality bar — so you can raise token volume without raising spend.
Model library
Frontier-class open models, tuned variants and small specialists — all served on the same runtime, all callable from the same endpoint.
CarraWorks-L 5.2
Frontier-class general reasoning
CarraWorks-K3
Long-context agentic workhorse
CarraSeek-V4-Pro
Reasoning with visible thinking traces
CarraQ3-A95B
Sparse MoE, high throughput per dollar
CarraMini M3
Fast, cheap, high-volume classification
CarraLightning 30B-A3B
Latency-optimized serving profile
CarraVision 12B
Documents, screenshots and charts
CarraEmbed 8B
Enterprise retrieval embeddings
Need a specialized reranker or embedding model trained on your data? Explore CarraZero →
Use cases
Code assistance
Low-latency completion and repo-aware review on models you control.
Conversational
Support, sales and internal assistants with streaming responses.
Agentic
Tool-calling agents with function schemas and long-running workflows.
Search
Embeddings, reranking and hybrid retrieval over your own corpus.
Multimodal
Vision, document and audio understanding on open VLMs.
Developer productivity
Batch pipelines, evals and internal tooling at team scale.
Start building in seconds, self-serve. Talk to us for enterprise deployments with faster speeds, lower costs and higher rate limits.
Standard
Best price per token for batch and background work.
From $0.10 / 1M input tokens
Priority
Lower queueing for interactive product traffic.
From $0.30 / 1M input tokens
Fast
Latency-first routing for real-time agents and voice.
From $0.60 / 1M input tokens
On-demand deployments
Pay per GPU second, billed by the second, for faster speeds, higher rate limits and lower cost at scale. Reserved capacity available on annual terms.
Embeddings
Training pricing
Supervised and preference fine-tuning are priced per 1M training tokens. Reinforcement fine-tuning jobs are priced per GPU hour, billed per second, at the same rate as on-demand deployments.
| Base model | LoRA SFT | LoRA DPO | Full-parameter SFT |
|---|---|---|---|
| Models up to 16B parameters | $0.50 | $1.00 | $2.00 |
| Models 16.1B – 80B | $3.00 | $6.00 | $6.00 |
| Models 80B – 300B | $6.00 | $12.00 | $12.00 |
| Models over 300B | $10.00 | $20.00 | — |
Training tokens can be estimated as tokens in your dataset × number of epochs. For tuning with intermediate thinking traces, multiply by the average number of conversation turns divided by two.
Including reasoning traces on assistant turns increases total tuned tokens, because multi-turn conversations are unrolled into user, assistant and thinking segments. Vision fine-tuning is billed per 1M tokens on the same table.
Why CarraWorks
One API, every model
OpenAI-compatible endpoints so you can swap models without rewriting your app.
Quality you can measure
Built-in evals, A/B routing and per-request tracing across models and tiers.
Private by default
Your prompts and datasets are never used to train shared models.
Enterprise controls
SSO, per-team quotas, audit logs, VPC and on-prem deployment options.
We'll benchmark your current provider against CarraWorks on latency, quality and cost — with your prompts, on your data.