Temper 1.5B is Qwen2.5-Coder-1.5B-Instruct fine-tuned on the same 10 thousand Rust examples as Temper 0.5B. Each example was written by DeepSeek-V4-Flash and passed the compiler and tests. There was no reinforcement learning at this stage, it goes into Temper 2.
On the standard HumanEval-Rust benchmark the model solves 50.5% of tasks on the first try. That is more than Qwen2.5-Coder-3B at twice the size and not far from the 7B and 9B models.
Fig. 1. Percentage of HumanEval-Rust tasks solved on the first try (pass@1, %, 156 tasks) by model size. The parameter scale is logarithmic. Settings follow the BigCode Models Leaderboard with code completion without a chat template, temperature 0.2, top_p 0.95 and 50 samples per task. Qwen and StarCoder2 models are the base versions.
Other Rust models
Among open models that other authors fine-tuned on Rust, Temper 1.5B is behind only the 7B and 14B models.
Fig. 2. Percentage of HumanEval-Rust tasks solved on the first try (pass@1, %) by open models fine-tuned on Rust.
Limitations
With ten attempts Temper 1.5B solves 67.0% of tasks, which nearly matches the result of Qwen2.5-Coder-3B (67.4%). The model has the same limitations as Temper 0.5B. It was trained only on Rust tasks with requests in English and without fill-in-the-middle mode, and the data also has few tasks on async code and web frameworks.