How it started
Despite using LLMs for most of the coding, there was always one thing I kept Googling: Bash commands. It's quite flow-breaking to pause work, open Google, type the full query, go to Stack Overflow or similar, and look up the syntax I wanted.
Intuitively, this always felt like something a small model would perform well at because you have a finite set of commands, well-defined syntax, and easy-to-generate training data. So one day I decided to actually find out. Over the course of the experiment, I tried six different models: SmolLM 135M, SmolLM 360M, Qwen3-0.6B, Qwen2.5-Coder-1.5B, Qwen3.5-0.8B, and Qwen3.5-2B. The 0.6B and 1.5B models showed the best performance, so those checkpoints were preserved.
The good stuff you probably want
I mentioned the experiment on Hacker News, and people asked for the models, training data, and a write-up.
- The fine-tuning dataset: 401,975 unique English request/Bash command pairs.
- The models: 1.5B GGUFs and 0.6B GGUFs.
- The CLI.
What started as a syntax reminder turned into an interesting exercise in teaching small models—and figuring out whether they were actually improving.
Astra ran the training loop
When I say I trained these models, I really mean Astra was the training controller, with different LLMs used to generate the synthetic dataset and another set of agents used to review and admit rows into it. Astra handled the training-and-evaluation workflow; I supplied the goal, constraints, and course corrections.
The starting point was coming up with a category of commands in a separate conversation and building a harness to generate and validate those commands. Initially we had about 30k samples. Then the loop was simple: train, evaluate (I got Astra to create some benchmarks), analyze, generate more data, brainstorm different ways to train, and run the next training runs. The whole process took about 10 days, most of it spent with models working on their specific tasks: generate, review, evaluate.
The prompt was some variation of (among other things):
We only care about English-to-Bash commands, nothing else. As long as we make gains in that domain, we are happy to sacrifice any abilities in any other domain. We want the best goddamn English-to-Bash model ever created in the history of humanity, do what you need to do.
The only real direction I provided was keeping a mental map of multiple trained checkpoints, prompting the models to compare those in specific categories, asking them to investigate why a checkpoint traded gains in one area for losses in another, and so on.
The released models use supervised fine-tuning with LoRA, starting from Qwen3-0.6B and Qwen2.5-Coder-1.5B-Instruct. The 4-bit versions surprisingly keep most of the gains of the 16-bit versions, and they are convenient to run on CPU systems.
More examples weren't always the answer
The published dataset contains 401,975 distinct request/answer pairs, covering file operations, text processing, Git, archives, networking, quoting, and composed commands. A row looks like this:
{"request":"list files in this directory","response":{"kind":"COMMAND","value":"ls"}}
Examples were synthetically generated, reviewed, and checked through execution where available. Not every row was individually executed. Different descriptions of the same command were intentionally retained.
One experiment made the difference between learning examples and learning behavior very clear. A 0.6B model improved from 8/72 to 64/72 on time-related training requests. On separate time-transfer requests, it stayed at 1/18. Meanwhile, routine-command performance fell from 70/102 to 60/102.
The model was learning the supplied answers. That wasn't the same as reliably applying them elsewhere.
Targeted repairs brought another problem: retention. A later 1.5B checkpoint passed five more ALFA-updated cases than the model I released, but lost 63 passes on the internal suite. I kept the more balanced checkpoint. Exporting also changed checkpoint rankings, so evaluating the actual quantized file became part of selection rather than an afterthought.
The detailed results preserve those comparisons, including the experiments that didn't help.
I ended up debugging the benchmark too
I used ALFA, a 300-task shell benchmark, alongside internal execution tests. Investigating disagreements uncovered both model mistakes and evaluator problems: argument transport, fixture resets, reference commands, and checks that could accept wrong outputs or reject correct alternatives.
That work became ALFA-updated. It keeps the task set but documents the repaired environment and grading rules. It's a separate benchmark definition, not a silent replacement for upstream ALFA.
Benchmark comparison
The table combines our local results with scores reported by the whatisit project and the barbarabhb model card. ALFA-updated? identifies the benchmark used: Yes for our repaired version, No for original ALFA. whatisit appears twice: its published original-ALFA score was 62.0%, while our ALFA-updated run measured 63.7%—a 1.7-percentage-point increase in the reported score.
| Model / configuration | Size | Pass rate | ALFA-updated? | Reported by |
|---|---|---|---|---|
| GPT-4o (published) | Cloud API | 73.0% | No | ALFA authors |
| EasyCommand 1.5B, Q4_K_M | 940.4 MiB | 212/300 — 70.7% | Yes | Our local run |
| nl2sh-3b | 1.9 GB | 65.7% | No | whatisit |
| barbarabhb / nl2sh-qwen25-coder-1.5b, Q4_K_M | 941 MB | 65.67% | No | barbarabhb |
| barbarabhb / nl2sh-qwen25-coder-1.5b, Q4_K_M + imatrix | 941 MB | 65.00% | No | barbarabhb |
| barbarabhb / nl2sh-qwen25-coder-1.5b, Q6_K | 1.2 GB | 64.33% | No | barbarabhb |
| barbarabhb / nl2sh-qwen25-coder-1.5b, Q8_0 | 1.6 GB | 63.67% | No | barbarabhb |
| whatisit / nl2sh-1.5b, Q4_K_M | 941 MB (reported) | 191/300 — 63.7% | Yes | Our local run |
| whatisit / nl2sh-1.5b, Q4_K_M | 941 MB | 62.0% | No | whatisit |
| Qwen2.5-Coder-7B, untuned | 4.4 GB | 61.3% | No | whatisit |
| EasyCommand 0.6B, Q8_0 | 767.5 MiB | 174/300 — 58.0% | Yes | Our local run |
| EasyCommand 0.6B, Q4_K_M | 461.8 MiB | 165/300 — 55.0% | Yes | Our local run |
| Qwen2.5-Coder-1.5B, untuned | 941 MB | 54.0% | No | whatisit |
21 more passes than whatisit (+7 points) on ALFA-updated, using each tool's own settings (256-token limit for EC, 64 for whatisit).
A specialist peaked at 217/300 (72.3%), but lost 63 internal-test passes. I shipped the more balanced checkpoint.
These tests informed development, not an independent holdout. GPT-4o's published 73–74% used original ALFA; similar scores don't establish parity. Full methodology and run details are in the evaluation notes.
Try it, or keep training
If you want to run it locally, ec embeds llama.cpp and
offers command preview and interactive execution. Start with --preview; a plausible
command can still be wrong. The target is GNU/Linux Bash, not every shell or platform.
If you want to experiment, the reusable pieces are:
- Training data: the full published deduplicated corpus, not just a sample.
- GGUFs: 0.6B and 1.5B.
- Trainable weights: 0.6B and 1.5B, including merged BF16 checkpoints and original LoRA adapters.
- Training notes and benchmark code.
Models and data are Apache-2.0; application and benchmark code are MIT. The flat dataset doesn't reproduce historical weighting, and it has no official test split. If you train with it, separate holdouts by task family rather than randomly splitting paraphrases.
I'd like to see what others can improve: better coverage, more reliable transfer, or fresh tests that expose something I missed. This started with wanting a command reminder. Sharing the materials seems like a good way to keep learning.