Skip to content
Supported training models

Supported training models

Choose a trainable base model and understand the current Rebyte Loop limits.

On this pageChoose a starting pointData limits

SDK examples on this page require the 0.5.0 source preview. npm publication is pending; use the HTTP endpoints directly in the meantime.

Rebyte Loop currently supports one trainable base model. Other models in the Agent API catalog can still produce work and feedback; training support is a separate capability.

FieldValue
ModelQwen3.6-35B-A3B
base_model when creating a model familyQwen/Qwen3.6-35B-A3B
Base inference model IDqwen3.6-35b-a3b
Training methodSupervised fine-tuning with low-rank adaptation (LoRA)
Training dataCaptured text messages and a reviewed assistant target
Reasoning modeDisabled for this training workflow

Preview inference is intended for evaluation and low-volume use. It does not include a production throughput SLA. Trained versions expire within seven days; the base inference model does not have a version expiration. Check each version's returned expires_at before use.

Choose a starting point

Create a Scenario-bound model family:

ts
const model = await rebyte.models.create({
  name: 'Order extraction',
  scenario_id: scenario.id,
  base_model: 'Qwen/Qwen3.6-35B-A3B',
}, { headers: { 'Idempotency-Key': 'order-model-001' } });

If a training run omits source, it starts from the family's published version, or from the base model if no version is published. You can also explicitly start from the base or from an owned ready version of the same family. The resolved choice is saved on the run and does not change on retries.

A base source exposes LoRA rank, seed and the attention, MLP and output-embedding training switches. It does not expose arbitrary layer numbers or individual projection names. Inspect the saved source and settings on each run before comparing experiments.

Data limits

Only completed Turns with supported captured text context and usable feedback are eligible. Positive feedback uses the original answer; a correction uses the complete corrected answer, even if the rating is negative. A negative rating alone is retained as feedback but is not a training target.

The current limits are 1,000 unique selected Turns, a 16 MiB frozen dataset, and 256 KiB per captured context plus target. Your run's batch and sequence-token limits also apply. Preparation fails on overlong examples; it does not truncate them. Tool/image context and historical Turns without capture are ineligible.

See training runs for selection, optimizer settings and progress. Check the version's returned expires_at, context limit, output limit and pricing before using it in your application.