Carraggon CarraWorks

The fastest way to run open models in production.

CarraWorks is Carraggon's inference and fine-tuning platform: serverless endpoints, dedicated GPU deployments, managed training and embeddings — one API, from first prototype to enterprise scale.

Time to first token

Seconds, no cold starts

Serving

Serverless or dedicated GPUs

Billing

Per token, or per GPU second

Deployment

Cloud, VPC or on-prem

Building with open models shouldn't mean building an inference team. CarraWorks handles the serving, the tuning and the scaling — you keep the weights, the data and the choice of model.

The platform

Six ways to ship.

Start serverless, move to dedicated capacity when volume justifies it, tune your own models, run your own RL loops, and route spend down with Nexus.

01

Serverless Inference

Get started in seconds with per-token pricing, zero setup and no cold starts.

Pay per token with high rate limits and postpaid billing. Standard, Priority and Fast tiers let you trade cost against latency per workload — same API, one line to switch.

  • No cold starts
  • High rate limits
  • Standard / Priority / Fast tiers
  • OpenAI-compatible API

02

On-Demand Deployments

Pay per GPU second for faster speeds, higher rate limits and lower cost at scale.

Dedicated capacity billed per second. Pin a model to your own GPUs for predictable latency, private networking and unlimited throughput once volume justifies it.

  • Billed per second
  • Dedicated GPUs
  • Predictable P99 latency
  • Autoscale to zero

03

Managed Training

Customize open models with your own data with minimal setup.

Supervised fine-tuning (SFT), preference tuning (DPO) and reinforcement fine-tuning, fully managed. Upload a dataset, pick a base model, and deploy the result as a LoRA add-on in minutes.

  • LoRA & full-parameter SFT
  • DPO preference tuning
  • Reinforcement fine-tuning
  • Deploy tuned models instantly

04

Embeddings & Rerankers

Retrieval quality is the ceiling on agent quality.

Hosted embedding and reranking endpoints priced per million input tokens, sized from compact 150M models to 8B-class encoders for the hardest enterprise corpora.

  • Compact to 8B encoders
  • Per-1M-token pricing
  • Batch & streaming
  • Pairs with CarraZero

05

RL Rollouts

The only dedicated rollout inference for teams running their own RL.

Rollout-optimized inference built for reinforcement learning loops: massive parallel sampling, in-flight weight sync and per-second GPU billing so your trainer never waits on the sampler.

  • Massive parallel sampling
  • In-flight weight updates
  • Bring your own RL harness
  • Billed per GPU second

06

Nexus

Increase token usage, reduce cost. An open-weight model layer behind your harness.

Nexus sits behind your existing agent harness and routes each call to the cheapest open-weight model that clears your quality bar — so you can raise token volume without raising spend.

  • Quality-aware routing
  • Drop-in behind your harness
  • Open-weight cost floor
  • Per-route spend reporting

Model library

The full CarraWorks model library.

Frontier-class open models, tuned variants and small specialists — all served on the same runtime, all callable from the same endpoint.

Text

CarraWorks-L 5.2

Frontier-class general reasoning

Agentic

CarraWorks-K3

Long-context agentic workhorse

Reasoning

CarraSeek-V4-Pro

Reasoning with visible thinking traces

MoE

CarraQ3-A95B

Sparse MoE, high throughput per dollar

Small

CarraMini M3

Fast, cheap, high-volume classification

Fast

CarraLightning 30B-A3B

Latency-optimized serving profile

Multimodal

CarraVision 12B

Documents, screenshots and charts

Embeddings

CarraEmbed 8B

Enterprise retrieval embeddings

Need a specialized reranker or embedding model trained on your data? Explore CarraZero →

Use cases

Built for the workloads teams actually run.

Code assistance

Low-latency completion and repo-aware review on models you control.

Conversational

Support, sales and internal assistants with streaming responses.

Agentic

Tool-calling agents with function schemas and long-running workflows.

Search

Embeddings, reranking and hybrid retrieval over your own corpus.

Multimodal

Vision, document and audio understanding on open VLMs.

Developer productivity

Batch pipelines, evals and internal tooling at team scale.

Pricing to seamlessly scale from idea to enterprise.

Start building in seconds, self-serve. Talk to us for enterprise deployments with faster speeds, lower costs and higher rate limits.

Standard

Best price per token for batch and background work.

From $0.10 / 1M input tokens

Priority

Lower queueing for interactive product traffic.

From $0.30 / 1M input tokens

Fast

Latency-first routing for real-time agents and voice.

From $0.60 / 1M input tokens

On-demand deployments

Pay per GPU second, billed by the second, for faster speeds, higher rate limits and lower cost at scale. Reserved capacity available on annual terms.

Embeddings

Up to 150M parameters$0.008 / 1M input tokens
150M – 350M parameters$0.016 / 1M input tokens
8B-class encoder$0.10 / 1M input tokens

Training pricing

Managed training, priced per million training tokens.

Supervised and preference fine-tuning are priced per 1M training tokens. Reinforcement fine-tuning jobs are priced per GPU hour, billed per second, at the same rate as on-demand deployments.

Base modelLoRA SFTLoRA DPOFull-parameter SFT
Models up to 16B parameters$0.50$1.00$2.00
Models 16.1B – 80B$3.00$6.00$6.00
Models 80B – 300B$6.00$12.00$12.00
Models over 300B$10.00$20.00

Training tokens can be estimated as tokens in your dataset × number of epochs. For tuning with intermediate thinking traces, multiply by the average number of conversation turns divided by two.

Including reasoning traces on assistant turns increases total tuned tokens, because multi-turn conversations are unrolled into user, assistant and thinking segments. Vision fine-tuning is billed per 1M tokens on the same table.

Why CarraWorks

Production-grade from the first request.

One API, every model

OpenAI-compatible endpoints so you can swap models without rewriting your app.

Quality you can measure

Built-in evals, A/B routing and per-request tracing across models and tiers.

Private by default

Your prompts and datasets are never used to train shared models.

Enterprise controls

SSO, per-team quotas, audit logs, VPC and on-prem deployment options.

Bring one workload. See it running today.

We'll benchmark your current provider against CarraWorks on latency, quality and cost — with your prompts, on your data.