This is the flow my current team and I have settled into, because it fits how we work. It is not a rulebook I carry everywhere. Every team has its own rhythm, and I am happy to learn a different one and adapt. The shape changes with the team, and I am glad when it does.
I do not have one workflow. I have a few, and in practice several are in motion at once rather than one after another. Validation produces evidence, the evidence drives a decision, the decision gets documented, and it has to reach the people already building on the system without breaking their work. They interlock and overlap, so no single one is the centre or the starting line. This page is the process itself. To see it applied to a real project, read the case studies.
Before a change ships, I sit with the people who actually use it, designers and developers, and watch where they get stuck. What confuses them is what I fix.
I work out the real problem, then decide with numbers I can check rather than taste.
I record the thinking behind a decision so others can follow the logic and decide for themselves, including changing it.
I get the change adopted without breaking the people already building on the system.
How I check that a new structure works with the people who actually use it before it ships. That means designers, but also developers, because in some teams developers work in Figma with no designer on hand.
Before anything else, I name the question the session has to answer, the assumption I am worried about, so the whole thing is built to prove or disprove something specific rather than to gather vague reactions.
I set real tasks, like browsing the tokens with no help, then rebuilding a reference card from a picture, and I build the mockups they will use.
I run them myself, in small batches, with a small panel that mixes designers and developers.
I score their choices against the real Figma bindings, not my memory of the call, so a misremembered moment does not become a reported result.
How someone navigates, and what they do not trust, tells me more than whether they got it right. The right-or-wrong count is the smaller signal.
Notes per person, the patterns across all of them, the findings ranked, and a clear list of what the developers need to pick up.
The findings go straight back into the naming and structure, and I check the direction against how mature systems solve the same problem.
How a foundation change actually happens: typography, spacing, colour.
I do not take a proposal at face value. I look for the underlying property that actually differs, because the reported symptom is usually not the real cause. A question framed as scope is often really about one measurable property underneath it.
Figma, the token files, Storybook, and the open threads. I read them together so I am seeing one picture, not several half-pictures that quietly disagree.
Contrast, line-height, and legibility floors are measured against the accessibility rules, not judged by eye. Judgment still does plenty of the work, but the contentious calls get backed by measurement.
I work through the options, say why I am dropping the ones I drop, and change course when the rendered result proves me wrong.
Because foundation work rewards getting it right over getting it fast, I build the change for real on the branch: the specimen, the same thing in context, and A/B versions. People can then react to the actual result, not an abstract argument.
I own the design part and hand the code-side items to developers rather than half-doing them, so each one lands with the person who can fix it well. One team, different parts of the same flow.
An ADR records the decision and the options I did not take, with the reason for each. So anyone coming later can see the logic I was working in and take it forward, or revisit it, from an informed place rather than a blank page.
Findings go on the relevant issue, kept to design only. I draft the comment first and post it once I am happy with it, and I leave open questions open instead of pretending they are settled.
Two sides of the same job: writing my own docs, and reviewing other people's.
Branch off the latest main. When I want to keep things tidy I work in a separate git worktree, so my own local files never end up in the branch by accident.
Only the current docs, never the frozen older snapshots, and I add any new page to the sidebar so it can actually be found.
British English, the same callout patterns as the rest of the docs, live Storybook examples, and no broken links.
I run the docs locally to look at them, then run the checks and the full build, which fails if a link is broken. A pre-push hook runs the same checks again as a safety net.
I add files by name, never everything at once, and check the list of changes before pushing so nothing sneaks in.
Marked as docs so it does not trigger a release. Green checks and one approval, then it merges.
I review with a focus, for example the interactive states, rather than trying to comment on everything at once.
The token CSS, the Figma bindings, the real contrast values. A claim that cannot be traced back to the source does not ship, whoever wrote it.
Fix it in place, or move it to its own page. For states, a simple table can make them much easier to read.
If I am proposing a change, I make it work first, so the suggestion is proven rather than a guess.
I keep a review focused on design, and pass code or pipeline points to the people who can fix them. One team, different parts of the same flow, so nothing gets dropped.
A decision only counts once the teams already building on the system can move to it without breaking their work.
The new version lands next to the current one rather than replacing it, so nothing in flight breaks the moment it merges.
The old version is marked as on its way out but kept working, so teams migrate on their own schedule instead of being forced through a breaking change.
Before anything is called stable, it is checked against the components and products that actually consume it, not just in isolation.
The change rolls out in stages rather than all at once, so a problem shows up small and early instead of everywhere at the same time.
The people building on the system get a clear path and clear timing, because a change no one understands is a change no one adopts.
The flows above run across a set of tools that have to stay in agreement. Here is each one and why it earns its place.
Where the design system lives and where I make foundation changes real: specimens, in-context frames, and A/B variants on a branch, so people react to the actual result rather than an abstract argument. I also drive Figma programmatically (plugin and MCP) for bulk work, like reorganising the UI kit, repairing components, and renaming variant properties across component sets during a migration.
Why: it is the single design source of truth, and where designers and developers actually pick tokens and components, so validation and changes have to happen here.
The bridge between Figma and code. I verify decisions against the real token values, not the design file's appearance.
Why: tokens are where the system working beyond the design file is won or lost. If Figma and the CSS disagree, the token layer is where I catch it.
The live component reference. I check that a decision holds in the built component, and embed examples directly in the docs.
Why: it closes the loop between what Figma says, what the tokens say, and what actually renders.
Where guidance is written and versioned. I edit the current version only, verify with a local build that fails on broken links, and register pages in the sidebar.
Why: documentation is how a decision survives past me, and where adoption starts.
Where decisions get made in the open and recorded. Findings go on the relevant issue, the reasoning and the rejected options go in an ADR, and changes ship through a reviewed PR.
Why: it keeps the trade-offs visible and lets anyone later see the logic and revisit it from an informed place.
Clean, single-purpose branches, and a separate worktree when I want parallel work without my personal files leaking in.
Why: every change should be small, reviewable, and traceable.
I measure contrast and line-height against the rules rather than judging by eye, and flag sizes too small to read.
Why: accessibility and trust are details you cannot eyeball. Numbers make the call defensible.
For multi-source audits, competitive research, Figma bulk operations, and drafting, always with evidence discipline: every fact traces back to a source, and hunches are never written up as results.
Why: it makes good design decisions easier to apply at scale, for people and tools alike, but only if the receipts are kept.
I look for what is actually wrong before I start solving it, because the reported symptom is rarely the real problem.
The team checks claims against the source (Figma bindings, token CSS, real contrast values) before making them. A hunch never gets written up as a result.
Some things belong in the system, some are edge cases better handled by one component. I decide which on purpose.
I keep the values real layouts need, even when a cleaner formula would drop them.
I record the option I rejected and why, so the trade-off stays visible for whoever comes next.
I own my part and hand work to whoever can do it best, not because it is "not my job", but because design intent, tokens, components, and docs are one continuous thing we own together.
I lean on AI for audits, research, and Figma work, but every fact still has to trace back to a source.
Nothing reaches a PR or a colleague until I have read it back to myself first.