draft is a single binary. Point it at one or more sources; each becomes its
own draft, processed as a queue.
draft [flags] [more-sources...]
Bare filenames resolve against ~/Drop/Drafts/Sources. PDF, Markdown and text
work everywhere; DOCX is macOS only.
Flags
Flag
Description
--engine <mode>
auto (default), ollama, or a provider name
--extract-engine <m>
Backend for claim extraction (default: --engine)
--write-engine <m>
Backend for writing the article (default: --engine)
--model <name>
Session-provider model override (e.g. opus)
--experimental
Let auto mode use experimental providers
--num-ctx <n>
Ollama context window (default 8192)
--num-predict <n>
Ollama max output tokens (default 6000)
--force-new
Draft even if today's folder already has one
--merge
Combine all sources into one draft
--resume
Reuse a verified claim ledger from an earlier attempt
--dry-run
Report what a run would do, without calling a model
--review <draft>
Enhance an existing draft with surgical edits
--frontmatter <f>
Regenerate frontmatter and final document
--keep-artifacts
Keep the claim ledger beside a successful draft
--print
Run without the TUI; print draft paths to stdout
--json
Run without the TUI; one JSON object per job
--completion <sh>
Print a completion script: bash, zsh, fish
--version
Print version and exit
Choosing a backend
In auto mode draft takes the first installed stable CLI on your PATH and
drives it through its own login. No token is read, stored or logged. If a call
fails — you are offline, or not logged in — the chain advances to the next
backend and stays there for the rest of the run. There is deliberately no
up-front network probe: a flaky connectivity check must never be what
downgrades an online machine to the local model.
Providers, in auto-selection order: claude, copilot, codex, agy,
cursor-agent, amp, crush, goose, grok, qwen. The last four are
experimental and skipped unless you pass --experimental.
Routing stages separately
Extraction is a dozen cheap, mechanical calls per paper. Writing is one call
that decides the article's quality. They do not need the same backend:
Every numeric variable is clamped at both ends. A value outside its range is
not silently ignored: the default is used and a warning goes to stderr, because
a tunable you believe took effect but did not is worse than one that was
rejected outright.
Article sets
One article, three files, always in step:
2026-07-30/
├── source/2026-07-30--body.md # the article — edit this
├── yaml/2026-07-30--frontmatter.yaml # adjacent frontmatter
└── final/2026-07-30--final.md # combined, ready to publish
Edit the body, then regenerate the other two in place with
draft --frontmatter <body.md>. Three rules make that safe to run at any time:
the filename is the article's identity, your curated fields always win, and
reprocessing unchanged input rewrites every file byte for byte identically.