Internal pilot
Persona-guided synthetic journey evaluation
AI-assisted pre-check · Reads a validated persona repository · Complements research with people, never replaces it
Agent task completionWhether the agent finished each task in the persona context
Simulated step countSteps taken versus the intended happy path
Detected friction eventsPoints where the agent stalled, backtracked or guessed
Journey completionWhether the end-to-end flow could be completed at all
Challenge. Usability testing stalls on setup, not analysis: picking the right persona, writing a task script, arranging a session. By the time results come back, the design has often moved on.
Action. The agent reads from the validated persona repository design and research already use, works out which persona a journey is built for, walks the flow in that context, and reports where it struggled — before anyone spends a research session on it.
Result. Obvious journey risks get caught in minutes, so moderated sessions with real participants are spent on the questions only people can answer.
Limits and honest caveats
- This produces synthetic evidence. An agent is not a user, and its output is not a usability finding.
- Metrics are named for what they are — agent task completion, not human task success — so they can't be mistaken for results from research with participants.
- Known weaknesses: model bias, persona fidelity, no emotional response, and no substitute for accessibility testing with disabled users.
- Accuracy against real moderated sessions has not been benchmarked yet; that comparison is the next step.
Live in team workflow
Heuristic Scoring Tool: heuristic evaluation as a repeatable, comparable score
Evaluation framework · Based on HFI's NPCI/VIMM model · 20 checks · Composite out of 100
Tier 1 — automated (~30%)Linters, contrast and accessibility audits produce a draft score
Tier 2 — AI-assisted (~45%)A vision-language model drafts a rating with rationale and evidence
Tier 3 — human (~25%)Checks needing task context are scored by a designer
Human sign-offEvery tier is a draft until a designer accepts or overrides it
Challenge. Heuristic evaluation usually produces an inconsistent one-off audit — different evaluators, different criteria, and no way to compare this quarter to last.
Action. Four focal points — Navigation, Presentation, Content and Interaction — with 5 checks each, rated 1–5, summing to a composite out of 100. Any check rated 1 or 2 generates a prioritised issue with evidence attached.
Result. The same 20 checks and the same scale every time, so a score can be tracked release over release and always carries its reasons.
Limits and honest caveats
- Automation never finalises a score. Tier 1 and Tier 2 produce drafts; a designer accepts or overrides each one, including the deterministic checks.
- Score bands are an internal convention for prioritising work, not an empirically calibrated threshold, and no band blocks a release on its own.
- The 80% agreement figure between AI drafts and human scores is a target being tracked, not a measured result.
- The rubric is adapted from Human Factors International's NPCI/VIMM model by a practitioner trained in it; naming and licensing would need HFI's agreement before any public product use.
Prototype
Craftline: design-system governance as an operating system
Governance model with an interactive prototype · Figma libraries in scope
Library auditRules rate violations Critical, High, Medium or Low
Drift detectionDetached instances, deprecated sources, local overrides
Component lifecycleIntake → Design Authority → Engineering → Release
Exception managementEvery exception has an owner, evidence and an expiry date
Challenge. Large organisations lose design-system consistency because dozens of teams, multiple libraries, local exceptions and accessibility requirements make drift inevitable without a system to catch it.
Action. Craftline scans the libraries in scope, flags what is actually broken, and routes new and changed components through four gated review stages with evidence attached.
Principle. Coaching, not blame — designers see remediation guidance, the design-system team can override and annotate any decision, and every decision stays auditable.
Limits and honest caveats
- Stage is a prototype plus a stakeholder presentation: the workflow is designed and demonstrable, not running against production files at organisation scale.
- Scope is the libraries connected to it, not "every Figma file".
- Designer scorecards carry a real risk of feeling like surveillance. They are scoped to coaching conversations, are explainable, and are not used for ranking or performance rating.
- Permissions, data retention, appeal route and the accountable governance owner are part of the model before any pilot.