The problem
A survey is a structured file: questions, answer types, translations, and rules that decide what each respondent sees next. AI agents are fast at editing files like that. What they can’t do is tell you reliably whether their edit broke something.
When a survey breaks, the person who finds out is the respondent. A broken rule can hide a question or leave someone stuck. A control can look fine and never record the answer. By the time anyone notices, the responses are already lost.
At the university where I work, I built a workflow to catch those problems before a change goes anywhere. This page describes the pattern; the examples are illustrative.
The context
The platform’s preview helped, but only up to a point. A file could import cleanly and still fail in the editor. A local rendering could look right and behave differently on the real platform. Seeing an answer selected on screen didn’t prove it would be stored.
Manual testing had a different problem: it only covers the path you happen to click through. Branching, required answers, language switches and mobile layouts create many paths through the same survey.
The failures I ran into shaped the design. Switching language on mobile could translate the questions but leave the navigation buttons in the original language. A local layout could ignore a mobile setting. A custom selection control could show a choice and still fail the platform’s required-answer validation. Each of those needs a different kind of check.
Architecture and decisions
Agents don’t edit the platform’s export format directly. They edit a structured representation where questions, their IDs and the relationships between them are explicit. Stable IDs matter, because skip logic points at questions by ID, and an edit that renames one silently breaks the logic.
The flow:
- The agent proposes a change to the structured source.
- A compiler produces the survey file.
- A separate checker inspects that final file: structure and supported types, translations, and the logic between questions.
- Only a file that passes, and has been reviewed, gets delivered.
I kept the checker separate from the compiler on purpose, so it can also check files that were produced or edited some other way. A successful compile on its own isn’t enough.
Reusable pieces, like the language selector, are installed by a deterministic installer rather than re-implemented by an agent each time. Installing one twice is detected, and conflicts are reported.
Review results are tied to the exact file through a content hash. If the file changes after review, the review no longer counts. That stops a pass from silently carrying over to a different version.
Proposal
ES: missing
Checks
- SchemaWaiting
- TranslationsWaiting
- Display and skip referencesWaiting
The list of known failures grew from real counterexamples. I compared changes against controls I knew worked, and kept the evidence types apart: import, editor, preview, and recorded answers. When an observation was inconclusive, I logged it as inconclusive instead of turning it into a rule.
What it does today
It’s a working, locally tested prototype covering structured authoring, compilation, independent checking and local review. Agents use it through an interface that provides authoring guidance, reusable components, diagnostics and the final survey file. Every delivered file carries its check results, its review and the rule version used.
Some parts, like the bilingual tooling and a few components, are already used internally. They prepare and validate translations without breaking the survey’s structure, and they’re reviewed on both desktop and mobile because the two can drift apart.
Integration isn’t finished, and I don’t have a measured time saving or adoption figure for this pattern yet.
Limits and what’s next
Matching the real platform’s behaviour is still partial. Rendering locally, importing, behaving correctly in the editor and recording answers are four different things to prove, and custom behaviour needs the most care.
The checks only catch what they encode. They can’t judge whether a question is well written or whether a translation means the right thing; that still needs a person.
Next: finish the integration, extend coverage with controlled counterexamples, and make each delivery report state plainly what was checked, which file was reviewed, and what still needs testing on the real platform.