Open call · № 04

Member of Technical Staff, Mid-training.

Instruction-tuned models are not directly suited to raw computer-use trajectories. Mid-training adapts the base model to our data before downstream RL.

You will design the SFT corpus, maintain the training recipe, build the eval loop, and decide when an initialization is ready for downstream training.

What you'll do
  • Domain adaptation to human screencasts.
  • Supervised fine-tuning on hindsight-enriched trajectories.
  • Catastrophic-forgetting mitigation.
  • Optimizer choices and data mixtures.