BetterForms
From the first question to an evidence-backed report.
- Status
- Live
- Link
- Visit
- Stack
- Next.jsTypescriptPythonStripe
A form that finishes the job after the last answer
BetterForms connects the whole loop of asking, answering, and learning. An author drafts a form, publishes it for a defined run, and shares a link. Respondents get one focused question at a time. Once the form closes, the author gets a report with suitable charts, the underlying answers, and qualified themes from written feedback.
The product hypothesis is simple: collecting answers is only half the work. The useful outcome is a report that helps someone understand what the answers say without hiding the evidence behind a summary.
01 · Ask
Build manually, use a template, or review an AI-generated draft.
02 · Answer
A focused, one-question flow with progress and review before submission.
03 · Learn
Close the form to read charts, written answers, themes, and a summary.
Give the author control before publishing
The editor supports manual creation, three starting templates, and AI drafting from a purpose or pasted questions. AI suggestions arrive as editable drafts; the author reviews and adopts them rather than publishing model output directly. The same question validation rules apply to manual and generated forms.
Forms can include short and long text, choices, multi-select, and ratings, with welcome and thank-you screens. The default path is linear. Branching is forward-only, and AI-generated branches require grounding in the author’s wording. One path calculation is reused across the player, validation, exports, integrations, and reporting so those surfaces agree on what each person actually saw.
Make responding focused, then report honestly
The player shows one question at a time, tracks progress, supports keyboard shortcuts for choices and ratings, saves a local draft, and offers a review screen before submission. If an earlier answer changes the branch, answers to now-unreachable questions are removed. A client-generated submission ID lets a retry converge on the same response.
Published form versions are immutable, and each response points to the version it answered. Report denominators include only people whose version contained a question and whose path reached it. Older questions and retired choices remain visible instead of being silently rewritten by later edits.
Charts follow the answer shape: ratings use ordered bars; multi-select results never use a donut because selections can add up to more than 100%; single choice uses a donut only when its distribution suits one. Every chart has a table view, and free text remains readable as individual answers.
Finding themes without treating clusters as facts
Written answers carry nuance but resist simple counting. Embeddings help locate related responses; a dense cluster alone cannot tell whether they describe the same recurring, actionable issue. BetterForms therefore uses a Python worker to propose groups, asks an LLM to review them against a theme policy, verifies the cited evidence, and counts membership in a separate pass.
This is a conceptual view of the worker’s stages. The portfolio page does not run the analysis; the live product runs it asynchronously after a form closes.
Conceptual pipeline
Open a stage to see what it does and produces.
1Embed
Represent each answer without losing its long-form detail
The Python worker splits answers that exceed the embedding model's token limit, embeds the pieces, and keeps their source response IDs. Questions are analysed independently; fewer than 20 text answers produce a too-few-responses result.
2Cluster
Find candidate neighbours
Density clustering deliberately proposes small groups of semantically similar answers. It is a discovery step, not the final theme count: sharing a topic does not necessarily mean describing the same issue.
3Review
Ask whether a group describes a recurring issue
An LLM reviews the question and examples from each candidate group. It can merge related groups, reject weak ones, or separate distinct issues mentioned in the same answer. Large jobs are reviewed in batches and then checked for duplicate themes across batches.
4Verify
Tie every published quote to an actual answer
Deterministic checks require each cited quote to appear verbatim in a supplied answer. A theme also needs support from at least three distinct responses and two distinct texts. Invalid evidence is removed; excessive citation errors reject the review.
5Assign
Count coverage in a separate pass
Final theme centroids come from validated evidence. Every remaining answer is compared against them and can support multiple themes, be marked borderline, or remain unthemed. Candidate cluster membership never becomes the published count by default.
Make slow or uncertain analysis a normal product state
Closing creates the ordinary report first. If text questions and allowance are available, a separate job starts theme analysis. The UI explains progress in plain language, and the app can reconcile missed worker callbacks or retry a failed run against the same closed response set. A separately generated executive summary uses report aggregates and selected theme evidence; its failure cannot erase the report.
The worker persists queued jobs and webhook events, checks request identity on repeat submissions, and signs callbacks to the app. The app validates the callback’s job and revision before storing a result. Rate and usage limits keep AI work bounded. Empty, inconclusive, too-few-responses, and failure outcomes have their own states; none need to be dressed up as an insight.
Fictional example report · fixed demo data
What could we do better?
26 written answers from a 48-response example form. These figures illustrate the report UI, not customer feedback or a live worker run.
Four praise or no-change answers remain unthemed; one answer is borderline. In the product, themes expand to show evidence quotes and the original answers remain available below.
What the evidence supports
Synthetic development cases helped test the pipeline. In one staged replay, pair F1 rose from 0.337 for over-split clustering to 0.688 after review and verified assignment. The replay reused recorded reviewer decisions, so those figures describe that development case, not production accuracy or customer impact.
The evaluation also exposed a harder problem: a supposedly theme-free fixture actually contained recurring issues, and its labels needed correction. The next meaningful test is a held-out set of human-labelled answers with a clear policy for theme boundaries. Quote provenance and recurrence checks make results auditable; they cannot, by themselves, guarantee that every proposed theme is useful.
