AI Lab

Building better ways to design

These are my experiments in AI and design operations. The fastest way to understand a new medium is to build with it — so each tool says what stage it is at, what evidence exists, and where it should not be trusted.

What the stages mean

  • Concept — defined, not built.
  • Prototype — working demo, not in day-to-day use.
  • Internal pilot — used on real projects by a small group, still being validated.
  • Live in team workflow — part of how work actually gets reviewed.
  • Public demo — anyone can try it.

Internal pilot

Persona-guided synthetic journey evaluation

Agent task completionWhether the agent finished each task in the persona context
Simulated step countSteps taken versus the intended happy path
Detected friction eventsPoints where the agent stalled, backtracked or guessed
Journey completionWhether the end-to-end flow could be completed at all

Challenge. Usability testing stalls on setup, not analysis: picking the right persona, writing a task script, arranging a session. By the time results come back, the design has often moved on.

Action. The agent reads from the validated persona repository design and research already use, works out which persona a journey is built for, walks the flow in that context, and reports where it struggled — before anyone spends a research session on it.

Result. Obvious journey risks get caught in minutes, so moderated sessions with real participants are spent on the questions only people can answer.

Limits and honest caveats

  • This produces synthetic evidence. An agent is not a user, and its output is not a usability finding.
  • Metrics are named for what they are — agent task completion, not human task success — so they can't be mistaken for results from research with participants.
  • Known weaknesses: model bias, persona fidelity, no emotional response, and no substitute for accessibility testing with disabled users.
  • Accuracy against real moderated sessions has not been benchmarked yet; that comparison is the next step.

Live in team workflow

Heuristic Scoring Tool: heuristic evaluation as a repeatable, comparable score

Tier 1 — automated (~30%)Linters, contrast and accessibility audits produce a draft score
Tier 2 — AI-assisted (~45%)A vision-language model drafts a rating with rationale and evidence
Tier 3 — human (~25%)Checks needing task context are scored by a designer
Human sign-offEvery tier is a draft until a designer accepts or overrides it

Challenge. Heuristic evaluation usually produces an inconsistent one-off audit — different evaluators, different criteria, and no way to compare this quarter to last.

Action. Four focal points — Navigation, Presentation, Content and Interaction — with 5 checks each, rated 1–5, summing to a composite out of 100. Any check rated 1 or 2 generates a prioritised issue with evidence attached.

Result. The same 20 checks and the same scale every time, so a score can be tracked release over release and always carries its reasons.

Limits and honest caveats

  • Automation never finalises a score. Tier 1 and Tier 2 produce drafts; a designer accepts or overrides each one, including the deterministic checks.
  • Score bands are an internal convention for prioritising work, not an empirically calibrated threshold, and no band blocks a release on its own.
  • The 80% agreement figure between AI drafts and human scores is a target being tracked, not a measured result.
  • The rubric is adapted from Human Factors International's NPCI/VIMM model by a practitioner trained in it; naming and licensing would need HFI's agreement before any public product use.

Prototype

Craftline: design-system governance as an operating system

Library auditRules rate violations Critical, High, Medium or Low
Drift detectionDetached instances, deprecated sources, local overrides
Component lifecycleIntake → Design Authority → Engineering → Release
Exception managementEvery exception has an owner, evidence and an expiry date

Challenge. Large organisations lose design-system consistency because dozens of teams, multiple libraries, local exceptions and accessibility requirements make drift inevitable without a system to catch it.

Action. Craftline scans the libraries in scope, flags what is actually broken, and routes new and changed components through four gated review stages with evidence attached.

Principle. Coaching, not blame — designers see remediation guidance, the design-system team can override and annotate any decision, and every decision stays auditable.

Limits and honest caveats

  • Stage is a prototype plus a stakeholder presentation: the workflow is designed and demonstrable, not running against production files at organisation scale.
  • Scope is the libraries connected to it, not "every Figma file".
  • Designer scorecards carry a real risk of feeling like surveillance. They are scoped to coaching conversations, are explainable, and are not used for ranking or performance rating.
  • Permissions, data retention, appeal route and the accountable governance owner are part of the model before any pilot.