Temper 1.5B

huggingface.co/temper-ai
October 2026

Temper 1.5B is Qwen2.5-Coder-1.5B-Instruct fine-tuned on the same 10 thousand Rust examples as Temper 0.5B. Each example was written by DeepSeek-V4-Flash and passed the compiler and tests. There was no reinforcement learning at this stage, it goes into Temper 2.

On the standard HumanEval-Rust benchmark the model solves 50.5% of tasks on the first try. That is more than Qwen2.5-Coder-3B at twice the size and not far from the 7B and 9B models.

Temper 1.5BTemper 0.5BQwen2.5-Coder-0.5BQwen2.5-Coder-1.5BQwen2.5-Coder-3BStarCoder2-3BQwen2.5-Coder-7BQwen3.5-9B
0 10 20 30 40 50 60 70 0.5B 1B 1.5B 3B 7B 9B parameters Temper 0.5B25.2 Qwen2.5-Coder-0.5B16.0 Qwen2.5-Coder-1.5B39.3 Qwen2.5-Coder-3B44.7 StarCoder2-3B25.1 Qwen2.5-Coder-7B57.3 Qwen3.5-9B56.7 Temper 1.5B50.5
0 10 20 30 40 50 60 70 0.5B 1.5B 3B 7B parameters Temper 0.5B25.2 Coder-0.5B16.0 Coder-1.5B39.3 Coder-3B44.7 StarCoder2-3B25.1 Coder-7B57.3 Qwen3.5-9B56.7 Temper 1.5B50.5
Fig. 1. Percentage of HumanEval-Rust tasks solved on the first try (pass@1, %, 156 tasks) by model size. The parameter scale is logarithmic. Settings follow the BigCode Models Leaderboard with code completion without a chat template, temperature 0.2, top_p 0.95 and 50 samples per task. Qwen and StarCoder2 models are the base versions.

Other Rust models

Among open models that other authors fine-tuned on Rust, Temper 1.5B is behind only the 7B and 14B models.

0 20 40 60 80 100 Strand-Rust-Coder-14B 60.0 Tessa-Rust-T1-7B 59.4 Temper 1.5B 50.5 Temper 0.5B 25.2 Mellum-4b-sft-rust 15.6
Fig. 2. Percentage of HumanEval-Rust tasks solved on the first try (pass@1, %) by open models fine-tuned on Rust.

Limitations

With ten attempts Temper 1.5B solves 67.0% of tasks, which nearly matches the result of Qwen2.5-Coder-3B (67.4%). The model has the same limitations as Temper 0.5B. It was trained only on Rust tasks with requests in English and without fill-in-the-middle mode, and the data also has few tasks on async code and web frameworks.