Train from reviewed Turns
Review Scenario data, freeze a training snapshot, and train a model without preparing token tensors.
On this page
Review and manage the dataStart a managed training runInitial defaultsAll data or untrained dataFollow, evaluate and publishHTTP referenceSDK examples on this page require the 0.5.0 source preview. npm publication is pending; use the HTTP endpoints directly in the meantime.
Your application can review business Turns, correct their answers, and explicitly start a training run. Rebyte freezes the selected examples, tokenizes them, runs supervised fine-tuning, and produces an immutable model version. You then evaluate that version and choose whether to publish it.
Rebyte Loop supports supervised fine-tuning for
Qwen3.6-35B-A3B. Use your organization API
key at https://api.rebyte.ai/v1. The organization must have training access;
inference uses normal model billing.
Review and manage the data
Bind business Sessions to a Scenario, then assess their completed Turns with feedback. Scenario Turn resources expose the input, final output, latest active feedback, and current training eligibility. These APIs can supply a data-review screen in your application.
import OpenAI from 'openai';
import { RebyteExtensions } from '@rebyteai/agent-extensions';
const rebyte = new RebyteExtensions(new OpenAI({
apiKey: process.env.REBYTE_API_KEY,
baseURL: 'https://api.rebyte.ai/v1',
maxRetries: 0,
}));
for await (const turn of rebyte.scenarios.turns.list(scenarioID, { status: 'completed' })) {
console.log(turn.id, turn.output_text, turn.latest_feedback, turn.training);
}
await rebyte.feedback.create({
turn_id: turnID,
rating: 'negative',
correction: '{"order_id":"A-102","refund_total":19.50}',
comment: 'The previous answer omitted the refunded item.',
}, { headers: { 'Idempotency-Key': 'review-order-A-102-001' } });
Use a correction as the complete desired answer, not an instruction such as
“fix the total.” Inspect training.eligible, training.reason, and
training.target before selecting a Turn. A negative rating without a usable
correction does not provide a supervised target.
The text-training preview uses exact context captured during execution. It does
not reconstruct past requests from today's Session settings. Old Turns without
that capture, tool/image context, and inputs beyond the supported limits are
ineligible. Review the reason; the server does not silently truncate or discard
unsupported context. Frozen examples expose the captured messages and desired
target, so you can inspect what the model will learn.
You can remove one assessment, clear a Turn's active feedback, or remove a terminal Turn from public data:
await rebyte.feedback.delete(feedbackID);
await rebyte.scenarios.turns.clearFeedback(scenarioID, turnID);
await rebyte.scenarios.turns.delete(scenarioID, turnID);
Turn deletion is soft deletion; unfinished Turns return 409. Deleted Turns
and their public history/Items become inaccessible. These operations affect
future selection. They do not rewrite examples already frozen in a training run
or remove historical context needed by an existing conversation. Later Turns in
the same Session become ineligible with context_contains_deleted_turn, because
their captured context may include the deleted text. Start a new Session for new
training data; the server does not guess whether compaction removed that context.
Start a managed training run
Create a model family bound to the Scenario, then submit a run with a stable retry key. The endpoint takes data-selection settings and prepares the text for training.
const family = await rebyte.models.create({
name: 'Order extraction',
base_model: 'Qwen/Qwen3.6-35B-A3B',
scenario_id: scenarioID,
}, { headers: { 'Idempotency-Key': 'orders-family-001' } });
// Persist this retry key and request body before sending the request.
const run = await rebyte.models.trainingRuns.create(family.id, {
data: { mode: 'all', turn_ids: [turnID] },
}, { idempotencyKey: 'orders-run-001' });
console.log(run.id, run.version_id, run.source, run.data, run.training);
console.log(run.example_count, run.dataset_sha256);
for await (const example of rebyte.models.trainingRuns.examples(family.id, run.id, { order: 'asc' })) {
console.log(example.turn_id, example.feedback_id, example.messages, example.target);
}
The returned ftr_ resource contains the persisted source, full settings,
selection, dataset hash, example count and preallocated ftv_ output ID.
Omitted settings are seeded once at admission and returned in full. Supply a
complete training object to override them; individual nested fields are not
partial updates. Optimizer, accumulation, batching and schedule controls are saved on the run.
When source is omitted, admission chooses the family's published version, or
the base model if none is published. You can explicitly choose a base source or
an owned ready source version. The resolved choice is frozen: a later publication
does not redirect an admitted run or a retry of it.
The same retry key and request return the original run and its original snapshot, even when feedback or publication has changed. A different request with that key conflicts. Keep the same key after transport failure. A terminal failed/cancelled run stays terminal; a new experiment needs a new key.
Initial defaults
If the family has no published version, the default source is:
{"kind":"base","rank":32,"seed":0,"trainMlp":true,"trainAttention":true,"trainUnembed":false}
The initial training settings are shown below. Every run returns its own full,
persisted values, so inspect the response when reproducing an experiment.
maxOptimizerSteps is a limit; it does not repeat a small dataset until the limit
is reached. Batches and accumulation groups determine the actual update count.
{
"epochs": 1,
"gradientAccumulationSteps": 4,
"finalPartialAccumulation": "flush",
"batchLimits": {"maxExamples": 4, "maxTokens": 32768, "maxSequenceTokens": 8192},
"lossNormalization": "weight_mean",
"adam": {"learningRate": 0.0001, "beta1": 0.9, "beta2": 0.95, "eps": 1e-8, "weightDecay": 0, "gradClipNorm": 1},
"schedule": {"kind": "constant", "warmupSteps": 0, "endLearningRate": 0.0001},
"checkpoint": {"everyOptimizerSteps": 10, "ttlSeconds": 86400},
"execution": {
"maxDurationSeconds": 3600, "rpcTimeoutSeconds": 60, "pollIntervalSeconds": 1,
"heartbeatIntervalSeconds": 5, "maxAttemptsPerRequest": 3,
"retryInitialIntervalSeconds": 1, "retryMaximumIntervalSeconds": 5,
"continueEveryGroups": 20, "pauseAtStart": false
},
"budget": {"maxOptimizerSteps": 128}
}
All data or untrained data
| Selection | Meaning |
|---|---|
mode: "all" | Default. Select eligible current data, including examples already present in the source's training ancestry. |
mode: "untrained" | Select eligible data whose sample hash is absent from the chosen source version's successful training ancestry. |
turn_ids: ["turn_…"] | Restrict selection to these Turns. An empty array is rejected. |
Omitted or null turn_ids | Consider every eligible Turn in the family's Scenario. |
“Untrained” is relative to the chosen source, not a global used flag. A run from
the base starts with empty ancestry. Training an unrelated branch does not
consume examples for this source. Failed or cancelled runs do not count as
successful ancestry. A source created through the low-level tokenized API has
no managed example lineage, so untrained cannot safely infer its data history
and is rejected with lineage_unavailable.
The current preview accepts at most 1,000 unique selected Turns and a 16 MiB frozen dataset. A captured context plus target is limited to 256 KiB. Explicitly selected Turns must all be eligible; an ineligible selection fails instead of quietly dropping entries. Token limits are checked during preparation, and an overlong example fails the run instead of being truncated.
The sample hash includes the Turn ID, captured messages and target. A feedback comment change alone does not make the same example new; a changed target does. The snapshot also records the feedback ID for review. Once admitted, each run retains its own snapshot. Later corrections, feedback clearing and Turn deletion do not change a running or completed job.
Follow, evaluate and publish
const state = await rebyte.models.trainingRuns.status(family.id, run.id);
console.log(state.run.status, state.run.preparation, state.workflow, state.model);
await rebyte.models.trainingRuns.control(family.id, run.id, { action: 'pause' });
// Control is asynchronous; retrieve status before assuming a pause has applied.
await rebyte.models.trainingRuns.control(family.id, run.id, { action: 'resume' });
Run states are queued, preparing, training, ready, failed and
cancelled. Preparation records tokenizer and batch statistics. Status also
exposes the underlying model/training workflow when available. Pause/resume
requests are recorded and forwarded to the training child when it is available;
check controlStatus and the child state to confirm application. Cancel stops admission of new work, but
work already in progress may finish. Pauses and retries count toward the
training deadline. stop_requested_at records a cancellation request while
the workflow waits for existing work and cleanup to settle; it does not by itself
mean the run has reached cancelled.
ready confirms that the produced version can serve inference. Test run.version_id
with held-out business tasks and review the feedback. Publishing is a separate
explicit operation:
await rebyte.models.publish(family.id, { version_id: run.version_id });
Publication affects new Sessions using the family alias. Existing Sessions keep their immutable version. Check the returned version expiry; publication does not extend it. Submitting feedback never starts training or publishes a model.
The training-run recipe provides separate commands:
node examples/agents-api/training-runs.mjs create MODEL_ID run-input.json RETRY_KEY
node examples/agents-api/training-runs.mjs examples MODEL_ID RUN_ID
node examples/agents-api/training-runs.mjs status MODEL_ID RUN_ID
node examples/agents-api/training-runs.mjs wait MODEL_ID RUN_ID
HTTP reference
All routes use organization ownership. Reads require tasks:read; writes require
tasks:write. They do not require an OpenAI-Beta header.
| Method and path | Purpose |
|---|---|
GET /v1/scenarios/{scenario_id}/turns | Review Turns, current feedback and eligibility; accepts optional status. |
GET /v1/scenarios/{scenario_id}/turns/{turn_id} | Retrieve one Turn's review data. |
DELETE /v1/feedback/{feedback_id} | Remove one assessment. |
DELETE /v1/scenarios/{scenario_id}/turns/{turn_id}/feedback | Clear active feedback; returns 204. |
DELETE /v1/scenarios/{scenario_id}/turns/{turn_id} | Soft-delete a terminal Turn. |
POST /v1/models/{model_id}/training-runs | Freeze data and admit training; requires Idempotency-Key. |
GET /v1/models/{model_id}/training-runs | List runs. |
GET /v1/models/{model_id}/training-runs/{run_id} | Retrieve persisted selection, settings, status and result ID. |
GET /v1/models/{model_id}/training-runs/{run_id}/examples | Read immutable snapshot examples. |
GET /v1/models/{model_id}/training-runs/{run_id}/status | { run, workflow, model }; workflow fields can be null. |
POST /v1/models/{model_id}/training-runs/{run_id}/control | Pause, resume or cancel; asynchronous 202. |
Lists use limit, order and after. Keep the same filters and order while
following cursors. Keep held-out evaluation Turns outside your training selection.