Quickstart

git clone https://github.com/LangerSword/smelt.git && cd smelt
uv sync
uv run smelt doctor
uv run smelt run examples/demo/run_length.py:rle_encode \
      --provider gemini --model gemma-4-26b-a4b-it
uv run smelt report runs/<run-id>

Verify and benchmark steps never need the network. Model runs need a provider key.

How it works

select ─▶ capture oracle ─▶ generate ─▶ build ─┬─ok─▶ verify ─┬─pass─▶ bench ─▶ accept ─▶ report
                          (model writes       │            │
                           lib.rs)            └─ errors ────┴─ mismatches ─▶ feed back ↺ (max 3 attempts)
                                                                             else reject (honest)

The orchestrator is deterministic; the model only acts inside generate. Transitions are checkpointed (state.json + events.jsonl), so runs resume.

Oracle

smelt takes ONE slow, pure Python function and freezes what it actually does into a differential oracle — 200+ recorded input/output cases, hashed and read-only. Code, not the model, checks that the Rust answers identically on every recorded case.

Safety

  • Scaffold + Cargo.toml are pinned (pyo3 =0.29.3) and supplied by smelt — the model fills in one function body only.
  • Static scan before build rejects: unsafe, extra crates, build.rs, std net/process/fs, include_*!, env!.
  • Builds happen in run-scratch dirs; the extension executes in a separate worker process under resource limits. smelt never writes into your project.

Benchmark rules

The speedup must clear a conservative 1.5x lower bound, not the median. Every number is labeled with the model that produced it. Nothing is rounded up.

CLI reference

Full generated reference on the Results page and in docs/cli.md.

smelt run file.py:fn [--provider p] [--model m] [--candidate lib.rs] [--cases N] [--json]
smelt capture file.py:fn [--cases N] [--json]
smelt resume run-dir [--json]
smelt report run-dir
smelt scan path [--json]
smelt apply run-dir          # VIEW ONLY
smelt doctor [--probe] [--json]
smelt evals [--provider p] [--model m] [--max-calls N] [--pause-s S] [--out-tag T]
smelt furnace

Exit codes: 0 accepted · 3 refused (impure) · 4 rejected · 5 harness error · 2 config.

Agent skill

skills/smelt/SKILL.md follows the Agent Skills open standard. Validated with agentskills validate. Install it into your agent so it can drive smelt itself.

Providers

ProviderModelStatus
geminigemma-4-26b-a4b-it (default), gemma-4-31b-itworking
localany llama.cpp / OpenAI-compatible modelworking (offline)
dogemma-4-31B-itwired, needs DIGITALOCEAN_TOKEN
openaiany OpenAI-compatible endpointwired
file / fakehand-provided / CI candidatesworking

Furnace TUI

Furnace is a Go TUI that replays finished runs and runs new ones live. It reads runs/<id>/events.jsonl: the loop as a horizontal flow, a results panel, a log pane, and the verdict — INGOT, SLAG, or ASH. It is a front end; the Python core stays the source of truth and it never runs generated code.

cd tui && go build -o furnace .    # Go 1.27+
./furnace
./furnace replay ../runs/<run-id>

Evals & meta-eval

The corpus is frozen and run sequentially. Every table is generated from real run output; each number is labeled with the model. The meta-eval is a false-accept gate: 13/13 hand-written wrong rewrites rejected. See Results.

Honest limits

  • Simple types only (str / int / float / bool / list[int] / list[str]); impure targets refused.
  • Integers tested inside the i64 range; big Python ints are outside the tested domain.
  • The oracle proves equivalence on the tested inputs — the claim is exactly that.
  • Small functions can legitimately fail to beat pyo3 call overhead — four corpus tasks were rejected at 1.22–1.46x.
  • Package layouts (relative imports) are not supported yet.

Security model

Generated Rust is treated as untrusted. The model fills one function body in a pinned scaffold. Static scan, isolated build, separate worker process under resource limits. Secrets resolve env → project .env → ~/.hermes/.env; never logged.

Contributing

Open a PR on GitHub. Keep the core deterministic; the model only acts inside generate. Every number must come from a real run and name its model.

License & AI use

MIT. See AI_USE.md for what was written by the build agent vs the human developer. Libraries and licenses: typer (MIT), rich (MIT), pyo3 (MIT OR Apache-2.0), maturin (MIT OR Apache-2.0), llama.cpp (MIT, local server). Site fonts (Space Grotesk, Inter, JetBrains Mono) are SIL OFL; GSAP is under the GreenSock No Charge license.