Results
Generated from evals/results.json at build time — never hand-edited. Failures and refusals stay in the table. Every number names the model that produced it.
Source: evals/results.json · provider gemini · model gemma-4-26b-a4b-it · ran on: laptop
| function | category | status | attempts | oracle cases | speedup | lower bound |
|---|---|---|---|---|---|---|
linear_score | float | rejected | 3 | 206 | 1.41x | 1.37x |
sigmoid_approx | float | accepted | 2 | 206 | 1.65x | 1.60x |
collatz_steps | int | accepted | 1 | 209 | 22.68x | 21.79x |
digital_root | int | accepted | 1 | 209 | 3.32x | 3.24x |
popcount | int | accepted | 1 | 209 | 4.99x | 4.72x |
reverse_digits | int | accepted | 1 | 209 | 2.19x | 2.16x |
sum_digits_sq | int | accepted | 1 | 209 | 3.01x | 2.83x |
triangular | int | accepted | 1 | 209 | 47.14x | 46.97x |
bucket_hash | list | accepted | 2 | 204 | 2.23x | 2.17x |
count_peaks | list | accepted | 1 | 204 | 2.13x | 2.02x |
longest_run | list | accepted | 2 | 204 | 1.64x | 1.54x |
max_subarray | list | accepted | 1 | 204 | 1.72x | 1.56x |
median_doubled | list | rejected | 3 | 204 | 1.38x | 1.35x |
prefix_max_sum | list | rejected | 3 | 204 | 1.46x | 1.41x |
sum_squares | list | accepted | 1 | 204 | 1.74x | 1.66x |
caesar_shift | string | accepted | 1 | 225 | 11.51x | 11.48x |
count_vowel_runs | string | accepted | 1 | 225 | 5.96x | 5.28x |
count_words | string | accepted | 1 | 225 | 7.36x | 6.74x |
is_palindrome_norm | string | accepted | 1 | 225 | 2.40x | 2.39x |
longest_alpha_run | string | accepted | 1 | 225 | 6.77x | 6.30x |
normalize_spaces | string | rejected | 3 | 225 | 1.22x | 1.20x |
reverse_words | string | accepted | 1 | 225 | 1.61x | 1.58x |
snake_to_camel | string | accepted | 1 | 225 | 2.66x | 2.49x |
top_word | string | accepted | 2 | 225 | 4.96x | 4.74x |
Speedups — measured, not modeled
The gold mark is the conservative lower bound — the acceptance rule reads it, not the median.
triangularcollatz_stepscaesar_shiftcount_wordslongest_alpha_runcount_vowel_runspopcounttop_worddigital_rootsum_digits_sqsnake_to_camelis_palindrome_normbucket_hashreverse_digitscount_peakssum_squaresmax_subarraysigmoid_approxlongest_runreverse_wordsFalse-accept gate
Wrong rewrites are rejected
13/13 deliberately wrong rewrites rejected by the meta-eval (plus a 16/16 tuple battery). The oracle catches wrong values, wrong exceptions, and panics — so a plausible-looking candidate cannot slip through.
Source: scripts/meta_eval.py · full corpus: evals/results.json
CLI reference
Commands, flags and exit codes. The help text is a captured snapshot of smelt --help; regenerate it with python3 site/scripts/gen_cli_ref.py --refresh.
Commands
smelt run <file.py:fn> [--provider p] [--model m] [--candidate lib.rs] [--cases N] [--json]
smelt capture <file.py:fn> [--cases N] [--json] # freeze an oracle, no generation
smelt resume <run-dir> [--json] # continue an interrupted run
smelt report <run-dir> # print the run report (markdown)
smelt scan <path> [--json] # candidates with tier + skip reasons
smelt apply <run-dir> # VIEW ONLY diff + fallback shim
smelt doctor [--probe] [--json] # environment + provider checks
smelt evals [--provider p] [--model m] [--max-calls N] [--pause-s S] [--out-tag T]
smelt furnace # launch the Go TUI if installed
Exit codes
| code | meaning |
|---|---|
| 0 | accepted — verified, speedup clears the conservative 1.5x lower bound |
| 2 | config error (bad target, dotted qualname, unsupported annotation on a required path) |
| 3 | refused — impure target or unsupported-but-understood shape (typed, with a hint) |
| 4 | rejected — no acceptable rewrite within budget (a correct, honest outcome) |
| 5 | harness/internal error (report the run dir) |
Live smelt --help
Usage: smelt [OPTIONS] COMMAND [ARGS]...
Smelt a slow pure Python function into a verified Rust extension. Only proven
metal ships. Bare `smelt` launches the Furnace TUI.
╭─ Options ────────────────────────────────────────────────────────────────────╮
│ --help Show this message and exit. │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─ Commands ───────────────────────────────────────────────────────────────────╮
│ run Run the full smelt loop on one function: path/to/file.py:function │
│ capture Capture + freeze the oracle only (no generation, no build). │
│ resume Resume an interrupted run from its checkpoint. │
│ report Print a run's report (markdown). │
│ scan List candidate functions under <path> (source for the TUI picker). │
│ apply Print the proposed change for an accepted run — VIEW ONLY, never │
│ writes. │
│ furnace Launch the Furnace TUI (Go binary) if installed; print build steps │
│ otherwise. │
│ doctor Check toolchain, sandbox, local server, and provider keys. Green = │
│ no FAIL rows. │
│ evals Run the frozen corpus sequentially; writes evals/results.json + │
│ results.md. │
╰──────────────────────────────────────────────────────────────────────────────╯