Models draft and review. Deterministic code decides what may be printed.
Wrong statistics can misdirect a real study, so a plan is never the output of one model. Here is the shape of the process, and why the combination works.
§1
Why this combination works
A language model is good at reading a study description the way a colleague would and drafting a complete, coherent plan. It is not a reliable judge of its own work. So at the steps where a wrong answer does the most harm, the design brief (the step the app calls Your design) and the analysis (the step it calls Your analysis and how many you need), the draft is read twice: once by Anthropic's strongest Claude model, and once, independently, by an OpenAI model, because a second company's model does not share the first one's blind spots. The two reviews are merged, worst case wins, and the drafter gets one corrective pass. Every other step goes without that model review. It gets a repair pass that refills any section the draft cut short, confined to the sections that step owns, and then the fixed checks described next, which every step gets.
Then the part that is not AI at all. Ask Mallard's own code runs a fixed set of statistical checks against the finished plan, the rules a biostatistician would apply by hand, and decides what may be printed. If the sample-size method does not fit what the study measures, the number is withheld rather than shown with a warning under it. No model can overrule that step. Models supply the reasoning; the code supplies the discipline.
§2
The design comes first, and it is free
The most consequential decision in a statistics plan is the study design, and a plan built on the wrong one is wrong throughout. So before any plan is drafted, Ask Mallard produces a design brief: the design it recommends and why, what the study would estimate, the conditions the design depends on, each with the record-level check that settles it, and the alternatives worth considering. Choosing an alternative rewrites the brief for that design, also free. A credit is spent only when you choose a design and generate the plan.
§3
What that gives you
Citations you can check. References are never written from memory. The model emits search queries; the app resolves each against PubMed live and shows what it found beside what was asked. A record that does not exist cannot appear in your plan.
Numbers that are tested. The power and sample-size engine uses the exact noncentral t distribution for comparisons of means and is validated against R, the standard statistical software, rather than against a previous run of itself.
Honesty where a number cannot be justified. Designs with no defensible closed-form calculation, a stepped wedge, competing risks, a noninferiority margin, a prediction model, a matched case-control, get the method by name and no figure.
A record of how the plan was made. Every plan carries the model that drafted it, whether the reviews raised and resolved anything, and a paste-ready disclosure sentence for your Methods section.
Each applied where it fits the design, with the methods literature behind every number cited in the plan itself.
§5
How it is tested
Eleven adversarial study descriptions, each written the way a clinician would describe the study rather than the method, generated repeatedly and scored against the rules they are designed to trip. Thirty checks in total. These checks are mechanical. They run on every plan, and a plan that fails one says so on the page rather than hiding it.
What this does not measure. The checks confirm that a plan named the right method. They cannot confirm the method was applied correctly, that its assumptions hold, or that the plan is good. Correct and good are different questions and only the first can be automated. A high pass rate is a floor under quality, never a ceiling.
Every number expires. Results are tied to a specific build on a specific date, and the underlying models change.
Rates are reported per engine. More demanding studies are routed to a stronger model, so a rate pooled across both is a rate for neither.
The evaluation is run by the person who built the tool, on scenarios that person chose. No independent evaluation has been carried out.
No outcome data. Nobody has shown that studies planned with this tool are better studies. That claim is not made here and should not be inferred from anything above.
§6
What it refuses to do
IRB determinations. A hedged orientation and the questions to take to your board. Only your IRB decides, and every export says so.
Final interpretation of results. Example results and discussion are labeled planning-stage structure with invented numbers.
A sample size it cannot justify.
Replacing your collaborators. Every plan grades its own complexity and says what that level needs.
Writing citations from memory.
A planning aid, not a substitute for a statistician or your IRB. Questions about the process: hello@askmallard.com.