Docs
Quickstart, how it works, CLI reference, providers, the Furnace TUI, evals, honest limits, and the security model.
Quickstart
git clone https://github.com/LangerSword/smelt.git && cd smelt
uv sync
uv run smelt doctor
uv run smelt run examples/demo/run_length.py:rle_encode \
--provider gemini --model gemma-4-26b-a4b-it
uv run smelt report runs/<run-id>
Verify and benchmark steps never need the network. Model runs need a provider key.
How it works
select ─▶ capture oracle ─▶ generate ─▶ build ─┬─ok─▶ verify ─┬─pass─▶ bench ─▶ accept ─▶ report
(model writes │ │
lib.rs) └─ errors ────┴─ mismatches ─▶ feed back ↺ (max 3 attempts)
else reject (honest)
The orchestrator is deterministic; the model only acts inside generate. Transitions are checkpointed (state.json + events.jsonl), so runs resume.
Oracle
smelt takes ONE slow, pure Python function and freezes what it actually does into a differential oracle — 200+ recorded input/output cases, hashed and read-only. Code, not the model, checks that the Rust answers identically on every recorded case.
Safety
- Scaffold +
Cargo.tomlare pinned (pyo3=0.29.3) and supplied by smelt — the model fills in one function body only. - Static scan before build rejects:
unsafe, extra crates,build.rs, std net/process/fs,include_*!,env!. - Builds happen in run-scratch dirs; the extension executes in a separate worker process under resource limits. smelt never writes into your project.
Benchmark rules
The speedup must clear a conservative 1.5x lower bound, not the median. Every number is labeled with the model that produced it. Nothing is rounded up.
CLI reference
Full generated reference on the Results page and in docs/cli.md.
smelt run file.py:fn [--provider p] [--model m] [--candidate lib.rs] [--cases N] [--json]
smelt capture file.py:fn [--cases N] [--json]
smelt resume run-dir [--json]
smelt report run-dir
smelt scan path [--json]
smelt apply run-dir # VIEW ONLY
smelt doctor [--probe] [--json]
smelt evals [--provider p] [--model m] [--max-calls N] [--pause-s S] [--out-tag T]
smelt furnace
Exit codes: 0 accepted · 3 refused (impure) · 4 rejected · 5 harness error · 2 config.
Agent skill
skills/smelt/SKILL.md follows the Agent Skills open standard. Validated with agentskills validate. Install it into your agent so it can drive smelt itself.
Providers
| Provider | Model | Status |
|---|---|---|
gemini | gemma-4-26b-a4b-it (default), gemma-4-31b-it | working |
local | any llama.cpp / OpenAI-compatible model | working (offline) |
do | gemma-4-31B-it | wired, needs DIGITALOCEAN_TOKEN |
openai | any OpenAI-compatible endpoint | wired |
file / fake | hand-provided / CI candidates | working |
Furnace TUI
Furnace is a Go TUI that replays finished runs and runs new ones live. It reads runs/<id>/events.jsonl: the loop as a horizontal flow, a results panel, a log pane, and the verdict — INGOT, SLAG, or ASH. It is a front end; the Python core stays the source of truth and it never runs generated code.
cd tui && go build -o furnace . # Go 1.27+
./furnace
./furnace replay ../runs/<run-id>
Evals & meta-eval
The corpus is frozen and run sequentially. Every table is generated from real run output; each number is labeled with the model. The meta-eval is a false-accept gate: 13/13 hand-written wrong rewrites rejected. See Results.
Honest limits
- Simple types only (str / int / float / bool / list[int] / list[str]); impure targets refused.
- Integers tested inside the i64 range; big Python ints are outside the tested domain.
- The oracle proves equivalence on the tested inputs — the claim is exactly that.
- Small functions can legitimately fail to beat pyo3 call overhead — four corpus tasks were rejected at 1.22–1.46x.
- Package layouts (relative imports) are not supported yet.
Security model
Generated Rust is treated as untrusted. The model fills one function body in a pinned scaffold. Static scan, isolated build, separate worker process under resource limits. Secrets resolve env → project .env → ~/.hermes/.env; never logged.
Contributing
Open a PR on GitHub. Keep the core deterministic; the model only acts inside generate. Every number must come from a real run and name its model.
License & AI use
MIT. See AI_USE.md for what was written by the build agent vs the human developer. Libraries and licenses: typer (MIT), rich (MIT), pyo3 (MIT OR Apache-2.0), maturin (MIT OR Apache-2.0), llama.cpp (MIT, local server). Site fonts (Space Grotesk, Inter, JetBrains Mono) are SIL OFL; GSAP is under the GreenSock No Charge license.