Member of Technical Staff, RL.
We use RL on computer-use tasks to train models for long-horizon work. The interface is streaming visual observation and low-level actions. Rewards come from task completion or rubric-based generative reward models.
You will design training procedures, build evals for failure modes, run experiments, and publish the results.
- •Training infrastructure.
- •Environment selection.
- •Algorithms and capability retention.
- •Evals and stability at long horizons.