Layer 03 · Models: see the whole stackRunix Models · Early access

Open models, tuned on your data

Runix Models fine-tunes small and open-weight models on your own domain data, measures them against the model you use today, and serves them on dedicated capacity behind the same OpenAI-compatible API as Runix Router.

Scoped and quoted before any training starts. Training data can come from your own team or from Runix Data.

engagement spec (example)
# agreed before any training starts
base_model:
  family: qwen      # licence checked
  size: small
method:
  tuning: lora      # then merged
  preference: dpo   # optional
data:
  from: runix-data  # or your own set
  held_out: true    # never trained on
evaluation:
  baseline: current-model
  suite: your-task
serving:
  mode: dedicated   # single-tenant
  api: openai-compatible
  via: runix-router

An illustrative spec, not a product API. Every field is agreed per engagement.

Open-weight families Qwen Llama Mistral Gemma DeepSeek How we choose a base model →

From base model to production endpoint

One engagement, six pieces of work. Each produces something you can inspect: a dataset, a training run, an evaluation against your current model, an endpoint.

Supervised fine-tuning

Training on examples of the task done well: records your team has labelled, or a dataset built with Runix Data. The dataset is treated as the first step of the work, not an afterthought.

LoRA and QLoRA

Parameter-efficient tuning trains a small set of added weights instead of the whole model, which keeps training affordable. QLoRA does the same over a quantised base model, and adapters can be merged into the weights for serving.

Preference tuning

Direct preference optimisation (DPO) on pairs of better and worse answers, for behaviour that is easier to judge than to write down: tone, format, and when to decline.

Evaluation before rollout

Task-specific test sets held out from training, and a before-and-after comparison against the model you run today. A tuned model is only better if you can show it to whoever signs off.

Dedicated deployment

Single-tenant serving sized to your real traffic, in your cloud account or on capacity arranged per engagement, with upgrades and node failures planned for. Quantised variants only where the evaluation shows quality holds.

One interface either way

The tuned model answers through an OpenAI-compatible API and can sit behind Runix Router next to hosted providers, so your application does not learn the difference.

Where it sits in the stack

Runix Models is the model layer. Training data comes up from Runix Data or from your own team, and the tuned model goes out through Runix Router, next to the hosted models your application already calls.

Your application

Product featuresAgentsInternal tools

L4 · Gateway

Runix RouterOne OpenAI-compatible endpoint for hosted models and yours

L3 · Models

Base modelOpen-weight, licence checkedTuningSFT, LoRA, QLoRA, DPOEvaluationAgainst your current modelDedicated servingSingle-tenant

L2 · Data

Runix DataDomain datasets, cleanedYour own dataExamples your team owns
Data moves up into training; requests move down through one endpoint. Your application calls the same API whether the answer comes from a hosted provider or from a model tuned for you.

What we tune and serve

Small and open-weight models, chosen for the task in front of them. We say which model we would start from, and why, before any work starts.

Open-weight families

Qwen, Llama, Mistral, Gemma, DeepSeek and other families whose weights are published under a licence that permits your use. Which one we start from depends on the task, the languages it has to handle and where it will run.

Licences, read first

Each family has its own terms. Some limit use above a number of monthly users, some require attribution or a naming convention for derived models, and some come with an acceptable-use policy. We read the licence against your use case before any work starts, and tell you if it does not fit.

Small and mid-size models

Models sized for one job done well: cheaper to serve on dedicated capacity, and practical to run inside your own account. Whether one is good enough for your task is what the evaluation is for, and we do not promise the answer in advance.

When a tuned model is the right call

Not every team needs its own model, and we will say so if you do not: a hosted model through Runix Router is often the cheaper answer. These are the cases where a model of your own earns its cost.

The data cannot leave

A regulator, a customer contract or an internal policy rules out sending it to a shared inference service. A model that runs in your own account keeps prompts and outputs where the data already is.

The volume is steady and large

At constant, high volume, dedicated capacity stops being the expensive option, and a smaller model needs less hardware to serve it.

The task is narrow and repeated

Classification, extraction, routing or one well-defined kind of answer, asked many times a day. A tuned small model may be enough; the evaluation decides, and we tell you when it is not.

The model is the moat

You have tuned something on proprietary data, and it should run on your terms rather than someone else's.

How early access works

Three steps, each with a person on the other end: Runix Models is an engagement, not a self-serve training console.

01Scope it and set a baseline

Tell us the task, the model you use for it today and the data you have. We agree how success is measured, and record your current model's results, before anything is trained. We reply within one business day.

02Tune and evaluate

We choose a base model with you, tune it, and test it on held-out examples against the baseline. You see the comparison and decide whether it ships; if it does not beat what you run today, we say so.

03Deploy, then hand over or run it

The model is served on dedicated capacity, in your cloud account or arranged per engagement, behind an OpenAI-compatible API. Your team takes it over, or we run and support it under a contract with Runix AI Inc.

Common questions

Which models can you tune?

Open-weight families whose licences permit your use, such as Qwen, Llama, Mistral, Gemma and DeepSeek. We check the licence against your use case before any work starts, and tell you if it does not fit.

What happens to the data we give you?

Data you give us for an engagement is used to train the model you asked for, and nothing else. Section 5 of the Terms of Service says so, and the engagement's statement of work sets out the details. Everywhere else, including Runix Router, your content is not used for training at all.

Who owns the tuned model?

That is written into the engagement contract before any training starts, together with where the weights are stored and who can access them. Nothing is trained until those terms are agreed.

Where does the model run?

In your own cloud account, next to your data, or on dedicated capacity arranged per engagement. Either way it is single-tenant and answers through an OpenAI-compatible API, so it can sit behind Runix Router next to the hosted models you already call.

Do we need Runix Data or Runix Router as well?

No. Bring your own training data, and call the model's endpoint directly if you do not use Runix Router. The layers connect, but neither is required.

How is it priced?

Per engagement, quoted after scoping and before any training starts. You see the scope and the number together, so there is nothing to reconcile afterwards.

Give a smaller model one job to own

Tell us the task, the model you use for it today and the constraint you cannot design around: residency, cost at volume or a model of your own. We reply within one business day, including when a tuned model is not the answer.

Request early access