Temper 0.5B

huggingface.co/temper-ai
October 2026

Temper 0.5B is a small model fine-tuned on top of Qwen2.5-Coder-0.5B-Instruct, with the emphasis on the Rust language. This allowed it to beat larger models, up to x3 in parameters. The goal was to test the capabilities of a tiny model on a narrow task.

First version

Initially Qwen3.6-35B-A3B was chosen as the teacher, and on an NVIDIA A40 48GB card it produced 2700 Rust examples in ~9 hours. The training data format was simpler than in the final version of Temper. Each example had a function with an interface, a doc comment, a request to complete it and an empty body with a single todo!(). The first attempt with this dataset already gave a noticeable gain compared with the base Qwen.

A task from the first version dataset

Task

Implement the Rust code below. Replace every todo!() with a working implementation.
Keep all type definitions and function signatures unchanged. Use only the Rust standard library.
Reply with the complete code in a single ```rust code block.

/// Calculates the number of full fuel tanks required and the amount of remaining fuel.
///
/// Given `total_fuel` and `tank_capacity`, returns a tuple `(full_tanks, remainder)`.
/// If `tank_capacity` is zero, returns `(0, total_fuel)`.
pub fn fuel_logistics(total_fuel: i64, tank_capacity: i64) -> (i64, i64) {
    todo!()
}
Teacher answer
/// Calculates the number of full fuel tanks required and the amount of remaining fuel.
///
/// Given `total_fuel` and `tank_capacity`, returns a tuple `(full_tanks, remainder)`.
/// If `tank_capacity` is zero, returns `(0, total_fuel)`.
pub fn fuel_logistics(total_fuel: i64, tank_capacity: i64) -> (i64, i64) {
    if tank_capacity == 0 {
        return (0, total_fuel);
    }
    let full_tanks = total_fuel / tank_capacity;
    let remainder = total_fuel % tank_capacity;
    (full_tanks, remainder)
}

After three epochs of training on these examples the first version solved 64.6% of trivial tasks on the first try versus 45.6% for the base Qwen2.5-Coder-0.5B. On easy tasks its result grew more than twofold, from 15.0% to 34.3%. The gain on the MultiPL-E sets turned out more modest and was 10 points on humaneval-rs and 7 on mbpp-rs.

Temper 0.5BFirst version of TemperQwen2.5-Coder-0.5B
0 20 40 60 80 100 trivial 81.8 easy 59.0 medium 18.0 humaneval-rs 34.7 mbpp-rs 38.8
Fig. 1. Share of tasks solved on the first try (pass@1, %). First version of Temper and the final Temper 0.5B against the base Qwen2.5-Coder-0.5B.

Second dataset

After that the dataset grew to 10 thousand examples, and DeepSeek-V4-Flash became the teacher. It is overall stronger than Qwen3.6-35B-A3B in programming and also cheaper via the API. The data became higher quality, duplicates and benchmark leaks were removed from it. Of the DeepSeek generations, ~60% of the examples were accepted, those that passed compilation and tests. Of them, ~65% were requests without code and ~35% were tasks with stubs. About 35% of the whole dataset were algorithmic examples.

A task without a stub

Task

Write a function that takes two sorted lists of strings and merges them into one sorted list, but only including strings that don't start with 'a' and don't end with 'z'.

Use only the Rust standard library.

Teacher answer
pub fn merge_filtered(list1: Vec<String>, list2: Vec<String>) -> Vec<String> {
    let mut result = Vec::with_capacity(list1.len() + list2.len());
    let mut i = 0;
    let mut j = 0;
    while i < list1.len() && j < list2.len() {
        if list1[i] <= list2[j] {
            let s = &list1[i];
            if !s.starts_with('a') && !s.ends_with('z') {
                result.push(s.clone());
            }
            i += 1;
        } else {
            let s = &list2[j];
            if !s.starts_with('a') && !s.ends_with('z') {
                result.push(s.clone());
            }
            j += 1;
        }
    }
    while i < list1.len() {
        let s = &list1[i];
        if !s.starts_with('a') && !s.ends_with('z') {
            result.push(s.clone());
        }
        i += 1;
    }
    while j < list2.len() {
        let s = &list2[j];
        if !s.starts_with('a') && !s.ends_with('z') {
            result.push(s.clone());
        }
        j += 1;
    }
    result
}

After training on this dataset Temper was further trained with reinforcement learning. The model wrote several solutions for each task, and a reward went to those of them that pass the tests. So it learned from its own answers, which were checked by the Rust compiler.

Results

Compared with the first version, this one gained noticeably on the Rust versions of the classic benchmarks from MultiPL-E. The share of tasks solved on the first try grew by 7.6 points on humaneval-rs and by 7.9 on mbpp-rs. With 95% probability the real gain lies between 3.0 and 12.3 points on the first set and between 5.1 and 10.7 on the second, so this is not chance.

0 10 20 30 40 50 60 70 80 90 100 0.5B 1B 1.5B 2B 3B parameters DeepSeek-V4-Flash, teacher, 284B (13B active)83.8 Qwen2.5-Coder-3B44.1 Qwen2.5-Coder-1.5B33.4 Qwen3.5-0.8B10.2 Qwen2.5-Coder-0.5B21.4 ThoxEdge-RustCoder-0.5B13.0 rust-mentor-0.6B8.1 Temper 0.5B37.5
0 10 20 30 40 50 60 70 80 90 100 0.5B 1B 2B 3B parameters DeepSeek-V4-Flash, teacher, 284B (13B active)83.8 Coder-3B44.1 Coder-1.5B33.4 Qwen3.5-0.8B10.2 Coder-0.5B21.4 ThoxEdge13.0 rust-mentor8.1 Temper 0.5B37.5
Fig. 2. Share of MultiPL-E tasks solved on the first try (pass@1, %, humaneval-rs and mbpp-rs together, 510 tasks) by model size. The parameter scale is logarithmic.

Our own benchmark

For Rust there are almost no benchmarks on which it makes sense to measure a tiny model. MultiPL-E translates HumanEval and MBPP tasks into Rust, but these are algorithmic functions over numbers and lists, originally written for Python. Structs with methods, traits, error handling through Result and work with crates almost never appear in them, although everyday Rust code consists of exactly this.

Therefore our own benchmark of 100 tasks at three difficulty levels was added to the two MultiPL-E sets. Each task consists of a request in plain text, a public interface and hidden tests. The model returns the full code, it is compiled together with the tests, and the task counts only if all tests pass. Tests the model writes itself do not affect the result. For each task it was checked that the reference solution passes all tests, a stub with todo!() compiles and passes none.

The requests in the benchmark are written in the same style as the Temper training data. That means the independent comparison with other models comes from the MultiPL-E sets, and our own tasks show how well the model copes with ordinary Rust code.

Other Rust models

Besides the Qwen models, the comparison includes two open models of the same size that other authors fine-tuned on Rust. ThoxEdge-RustCoder-0.5B from THOX.ai is built on the same base as Temper and trained on the open Strandset-Rust-v1 dataset. rust-mentor-0.6b is obtained from Qwen3-0.6B through QLoRA on the same dataset and intended as an assistant for learning Rust and code review. Its LoRA adapter was merged into the original Qwen3-0.6B before the measurement. Both models were tested under the same conditions as the rest, without a system prompt and without reasoning.

Table 1. Solved on the first try (pass@1). Small print shows the share of tasks solved in at least one of 10 attempts (pass@10).

Temper
0.5B
ThoxEdge-
RustCoder 0.5B
rust-mentor
0.6B
Qwen2.5-Coder
1.5B
Qwen3.5
0.8B
Qwen2.5-Coder
0.5B
Qwen2.5-Coder
3Bfor reference
DeepSeek-V4
Flashteacher, 284B
Trivial taskstrivial¹81.8%pass@10 94.0%55.8%pass@10 88.0%35.4%pass@10 68.0%70.0%pass@10 94.0%45.2%pass@10 84.0%45.6%pass@10 88.0%77.0%pass@10 96.0%97.8%pass@10 100.0%
Easy taskseasy¹59.0%pass@10 76.7%13.0%pass@10 50.0%13.7%pass@10 30.0%42.7%pass@10 73.3%13.3%pass@10 36.7%15.0%pass@10 50.0%51.0%pass@10 83.3%97.3%pass@10 100.0%
Medium tasksmedium¹18.0%pass@10 35.0%1.5%pass@10 10.0%0.0%pass@10 0.0%11.0%pass@10 45.0%3.0%pass@10 5.0%1.0%pass@10 5.0%25.0%pass@10 50.0%79.5%pass@10 100.0%
Function generationHumanEval-RS34.7%pass@10 57.1%13.3%pass@10 41.7%4.6%pass@10 16.7%31.0%pass@10 69.2%6.9%pass@10 20.5%16.9%pass@10 42.3%50.6%pass@10 87.2%89.6%pass@10 98.1%
Basic tasksMBPP-RS38.8%pass@10 54.8%12.8%pass@10 44.4%9.7%pass@10 27.1%34.4%pass@10 64.7%11.6%pass@10 33.3%23.4%pass@10 50.8%41.3%pass@10 72.9%81.2%pass@10 89.8%

Highlighted is the best pass@1 among models up to 1.5B. The ▲ sign marks 3B where it is higher than all models up to 1.5B. All Qwen models are taken in their Instruct versions. Measurement: temperature 0.8, top_p 0.95, 10 answers per task, pass@k by the unbiased estimator from the Codex paper. ¹ Internal set of Rust tasks.

Limitations and plans

Temper has noticeable limitations. Among models of comparable size the model is the best at solving a task on the first try, but with ten attempts its advantage disappears. On pass@10 Qwen2.5-Coder-1.5B beats it on medium, HumanEval-RS and MBPP-RS, and on the last two the difference reaches 10–12 points. Partly this is the price of reinforcement learning, which made the model answers more confident and more uniform. Qwen2.5-Coder-3B remains stronger than Temper on medium tasks and on both MultiPL-E sets. The benchmark own tasks are written in the style of the Temper training data, so only the MultiPL-E sets give an independent estimate.

The dataset will grow, it currently lacks tasks on web frameworks and asynchronous code. For autocomplete in the editor it is still to be decided whether the model needs fill-in-the-middle (FIM) mode.