Temper 0.5B
Temper 0.5B is a small model fine-tuned on top of Qwen2.5-Coder-0.5B-Instruct, with the emphasis on the Rust language. This allowed it to beat larger models, up to x3 in parameters. The goal was to test the capabilities of a tiny model on a narrow task.
First version
Initially Qwen3.6-35B-A3B was chosen as the teacher, and on an NVIDIA A40 48GB card it produced 2700 Rust examples in ~9 hours. The training data format was simpler than in the final version of Temper. Each example had a function with an interface, a doc comment, a request to complete it and an empty body with a single todo!(). The first attempt with this dataset already gave a noticeable gain compared with the base Qwen.
A task from the first version dataset
Implement the Rust code below. Replace every todo!() with a working implementation.
Keep all type definitions and function signatures unchanged. Use only the Rust standard library.
Reply with the complete code in a single ```rust code block.
/// Calculates the number of full fuel tanks required and the amount of remaining fuel.
///
/// Given `total_fuel` and `tank_capacity`, returns a tuple `(full_tanks, remainder)`.
/// If `tank_capacity` is zero, returns `(0, total_fuel)`.
pub fn fuel_logistics(total_fuel: i64, tank_capacity: i64) -> (i64, i64) {
todo!()
}
Teacher answer
/// Calculates the number of full fuel tanks required and the amount of remaining fuel.
///
/// Given `total_fuel` and `tank_capacity`, returns a tuple `(full_tanks, remainder)`.
/// If `tank_capacity` is zero, returns `(0, total_fuel)`.
pub fn fuel_logistics(total_fuel: i64, tank_capacity: i64) -> (i64, i64) {
if tank_capacity == 0 {
return (0, total_fuel);
}
let full_tanks = total_fuel / tank_capacity;
let remainder = total_fuel % tank_capacity;
(full_tanks, remainder)
}
After three epochs of training on these examples the first version solved 64.6% of trivial tasks on the first try versus 45.6% for the base Qwen2.5-Coder-0.5B. On easy tasks its result grew more than twofold, from 15.0% to 34.3%. The gain on the MultiPL-E sets turned out more modest and was 10 points on humaneval-rs and 7 on mbpp-rs.
Second dataset
After that the dataset grew to 10 thousand examples, and DeepSeek-V4-Flash became the teacher. It is overall stronger than Qwen3.6-35B-A3B in programming and also cheaper via the API. The data became higher quality, duplicates and benchmark leaks were removed from it. Of the DeepSeek generations, ~60% of the examples were accepted, those that passed compilation and tests. Of them, ~65% were requests without code and ~35% were tasks with stubs. About 35% of the whole dataset were algorithmic examples.
A task without a stub
Write a function that takes two sorted lists of strings and merges them into one sorted list, but only including strings that don't start with 'a' and don't end with 'z'.
Use only the Rust standard library.
Teacher answer
pub fn merge_filtered(list1: Vec<String>, list2: Vec<String>) -> Vec<String> {
let mut result = Vec::with_capacity(list1.len() + list2.len());
let mut i = 0;
let mut j = 0;
while i < list1.len() && j < list2.len() {
if list1[i] <= list2[j] {
let s = &list1[i];
if !s.starts_with('a') && !s.ends_with('z') {
result.push(s.clone());
}
i += 1;
} else {
let s = &list2[j];
if !s.starts_with('a') && !s.ends_with('z') {
result.push(s.clone());
}
j += 1;
}
}
while i < list1.len() {
let s = &list1[i];
if !s.starts_with('a') && !s.ends_with('z') {
result.push(s.clone());
}
i += 1;
}
while j < list2.len() {
let s = &list2[j];
if !s.starts_with('a') && !s.ends_with('z') {
result.push(s.clone());
}
j += 1;
}
result
}
After training on this dataset Temper was further trained with reinforcement learning. The model wrote several solutions for each task, and a reward went to those of them that pass the tests. So it learned from its own answers, which were checked by the Rust compiler.
Results
Compared with the first version, this one gained noticeably on the Rust versions of the classic benchmarks from MultiPL-E. The share of tasks solved on the first try grew by 7.6 points on humaneval-rs and by 7.9 on mbpp-rs. With 95% probability the real gain lies between 3.0 and 12.3 points on the first set and between 5.1 and 10.7 on the second, so this is not chance.
Our own benchmark
For Rust there are almost no benchmarks on which it makes sense to measure a tiny model. MultiPL-E translates HumanEval and MBPP tasks into Rust, but these are algorithmic functions over numbers and lists, originally written for Python. Structs with methods, traits, error handling through Result and work with crates almost never appear in them, although everyday Rust code consists of exactly this.
Therefore our own benchmark of 100 tasks at three difficulty levels was added to the two MultiPL-E sets. Each task consists of a request in plain text, a public interface and hidden tests. The model returns the full code, it is compiled together with the tests, and the task counts only if all tests pass. Tests the model writes itself do not affect the result. For each task it was checked that the reference solution passes all tests, a stub with todo!() compiles and passes none.
The requests in the benchmark are written in the same style as the Temper training data. That means the independent comparison with other models comes from the MultiPL-E sets, and our own tasks show how well the model copes with ordinary Rust code.
Write a function that returns how many characters are in a string. It should count each character individually, so it works the same whether the string is empty, plain English, or contains non-English letters.
Use only the Rust standard library. Use exactly this interface:
pub fn char_count(s: &str) -> usize;
Temper 0.5B answer
Solved in 8 attempts out of 10. Below is one of the solved ones.
pub fn char_count(s: &str) -> usize {
s.chars().count()
}
Make a bank account thing with a balance you can check, add money to, and take money out. Taking money out should fail with exactly "insufficient funds" if there isn't enough, and in that case the balance shouldn't change.
Use only the Rust standard library. Use exactly this interface:
pub struct Account { /* private fields */ }
impl Account {
pub fn new(balance: u64) -> Self;
pub fn balance(&self) -> u64;
pub fn deposit(&mut self, amount: u64);
pub fn withdraw(&mut self, amount: u64) -> Result<(), String>;
}
Temper 0.5B answer
Solved in 9 attempts out of 10. Below is one of the solved ones.
pub struct Account {
balance: u64,
}
impl Account {
pub fn new(balance: u64) -> Self {
Account { balance }
}
pub fn balance(&self) -> u64 {
self.balance
}
pub fn deposit(&mut self, amount: u64) {
self.balance += amount;
}
pub fn withdraw(&mut self, amount: u64) -> Result<(), String> {
if amount > self.balance {
return Err("insufficient funds".to_string());
}
self.balance -= amount;
Ok(())
}
}
Write a tree type with a value and a list of children, where children can be shared (the same child appearing in more than one place). Add two functions: one that returns the sum of the values of a node and everything below it, and one that adds a number to the value of a node and everything below it. A shared subtree should be visited and counted once for each place it appears.
Use only the Rust standard library. Use exactly this interface:
use std::cell::RefCell;
use std::rc::Rc;
pub struct Node {
pub value: i32,
pub children: Vec<Rc<RefCell<Node>>>,
}
pub fn tree_sum(node: &Rc<RefCell<Node>>) -> i32;
pub fn add_to_all(node: &Rc<RefCell<Node>>, delta: i32);
Temper 0.5B answer
Solved in 7 attempts out of 10. The example is chosen as illustrative and is not typical for the level, on average Temper solves 18% of medium tasks on the first try.
use std::cell::RefCell;
use std::rc::Rc;
pub struct Node {
pub value: i32,
pub children: Vec<Rc<RefCell<Node>>>,
}
pub fn tree_sum(node: &Rc<RefCell<Node>>) -> i32 {
let mut sum = node.borrow().value;
for child in node.borrow().children.iter() {
sum += tree_sum(child);
}
sum
}
pub fn add_to_all(node: &Rc<RefCell<Node>>, delta: i32) {
let mut current = node.borrow_mut();
for child in current.children.iter_mut() {
add_to_all(child, delta);
}
current.value += delta;
}
Other Rust models
Besides the Qwen models, the comparison includes two open models of the same size that other authors fine-tuned on Rust. ThoxEdge-RustCoder-0.5B from THOX.ai is built on the same base as Temper and trained on the open Strandset-Rust-v1 dataset. rust-mentor-0.6b is obtained from Qwen3-0.6B through QLoRA on the same dataset and intended as an assistant for learning Rust and code review. Its LoRA adapter was merged into the original Qwen3-0.6B before the measurement. Both models were tested under the same conditions as the rest, without a system prompt and without reasoning.
Table 1. Solved on the first try (pass@1). Small print shows the share of tasks solved in at least one of 10 attempts (pass@10).
| Temper 0.5B | ThoxEdge- RustCoder 0.5B | rust-mentor 0.6B | Qwen2.5-Coder 1.5B | Qwen3.5 0.8B | Qwen2.5-Coder 0.5B | Qwen2.5-Coder 3Bfor reference | DeepSeek-V4 Flashteacher, 284B | |
|---|---|---|---|---|---|---|---|---|
| Trivial taskstrivial¹ | 81.8%pass@10 94.0% | 55.8%pass@10 88.0% | 35.4%pass@10 68.0% | 70.0%pass@10 94.0% | 45.2%pass@10 84.0% | 45.6%pass@10 88.0% | 77.0%pass@10 96.0% | 97.8%pass@10 100.0% |
| Easy taskseasy¹ | 59.0%pass@10 76.7% | 13.0%pass@10 50.0% | 13.7%pass@10 30.0% | 42.7%pass@10 73.3% | 13.3%pass@10 36.7% | 15.0%pass@10 50.0% | 51.0%pass@10 83.3% | 97.3%pass@10 100.0% |
| Medium tasksmedium¹ | 18.0%pass@10 35.0% | 1.5%pass@10 10.0% | 0.0%pass@10 0.0% | 11.0%pass@10 45.0% | 3.0%pass@10 5.0% | 1.0%pass@10 5.0% | 25.0%pass@10 50.0% | 79.5%pass@10 100.0% |
| Function generationHumanEval-RS | 34.7%pass@10 57.1% | 13.3%pass@10 41.7% | 4.6%pass@10 16.7% | 31.0%pass@10 69.2% | 6.9%pass@10 20.5% | 16.9%pass@10 42.3% | 50.6%pass@10 87.2% | 89.6%pass@10 98.1% |
| Basic tasksMBPP-RS | 38.8%pass@10 54.8% | 12.8%pass@10 44.4% | 9.7%pass@10 27.1% | 34.4%pass@10 64.7% | 11.6%pass@10 33.3% | 23.4%pass@10 50.8% | 41.3%pass@10 72.9% | 81.2%pass@10 89.8% |
Highlighted is the best pass@1 among models up to 1.5B. The ▲ sign marks 3B where it is higher than all models up to 1.5B. All Qwen models are taken in their Instruct versions. Measurement: temperature 0.8, top_p 0.95, 10 answers per task, pass@k by the unbiased estimator from the Codex paper. ¹ Internal set of Rust tasks.
Limitations and plans
Temper has noticeable limitations. Among models of comparable size the model is the best at solving a task on the first try, but with ten attempts its advantage disappears. On pass@10 Qwen2.5-Coder-1.5B beats it on medium, HumanEval-RS and MBPP-RS, and on the last two the difference reaches 10–12 points. Partly this is the price of reinforcement learning, which made the model answers more confident and more uniform. Qwen2.5-Coder-3B remains stronger than Temper on medium tasks and on both MultiPL-E sets. The benchmark own tasks are written in the style of the Temper training data, so only the MultiPL-E sets give an independent estimate.
The dataset will grow, it currently lacks tasks on web frameworks and asynchronous code. For autocomplete in the editor it is still to be decided whether the model needs fill-in-the-middle (FIM) mode.