The problem
The organisation runs 50+ applications and products, each with different scope and maturity. The design system had grown organically across all of them, but not evenly. Teams built local components when the system didn't cover their case. Nobody had a clear picture of which patterns were proliferating, where, or whether the system needed new variants or entirely new components.
Only 14 of 32 components had been built. Five of the most-used, each appearing in 70%+ of screens, had no design system implementation at all. The component roadmap was reactive: build whatever the loudest team asked for next.
Engineering was blocked on components that looked complete in Figma but lacked the variants production actually used. Adoption stalled. Without a baseline inventory, every prioritisation decision was guesswork.
The objective was twofold. First, produce a single source of truth covering every component, pattern, and gap, severity-tagged so leadership could prioritise. Second, codify the methodology so the same audit could run again without me in the loop.
Approach
Auditing 50+ products was not realistic. As a team we chose to focus on the 12 applications with the most active design and development. The products where the design system mattered most day-to-day. That gave us 31 production screens across dashboards, planners, file managers, form editors, map interfaces, and calendar grids. Each screen averaged around 50 component instances. The most complex editors reached 140+.
The audit had two layers. First, a broad inventory across all screens to count and classify every component, tag severity, and surface the coverage gap. Second, a deep pattern analysis on one component, Card, to define and prove the full methodology: anatomy diagrams, cross-pattern matrices, variant gap analysis, and a component API recommendation.
I structured the methodology as seven sequential phases. Each phase produces an artifact the next one depends on.
- Inventory Every visible UI element, categorised into eight canonical buckets: Form Inputs, Buttons and Actions, Data Display, Navigation, Feedback and Messaging, Layout and Structure, Date and Time, Surfaces.
- Pattern identification For the target component, identifying distinct patterns. Two instances count as the same pattern only if they share layout, content slots, and interaction behaviour. Pure content differences do not qualify.
- Cross-pattern matrices Three matrices (Visual Properties, Interactivity, Content Slots) that make divergence visible at a glance.
- Variant gap analysis Observed variants minus documented variants equals gaps. Documented minus observed equals dead weight worth flagging.
- Component recommendation A clear API decision. Compound or single? Which variants earn their place? What should not be built in?
- Severity tagging Every finding gets a P0 to P3 tag using fixed thresholds based on instance count, app spread, and workaround availability.
- Report assembly Everything lands in a structured Markdown deliverable with a stable section order so downstream readers know exactly where to look.
The broad inventory (phases 1, 6, 7) covered all 32 components. The deep pattern work (phases 2-5) was applied to Card to define the methodology, producing the anatomy diagrams and matrices that became the most-used artifacts of the whole audit.
Findings
The audit surfaced the gap between perceived completeness and actual coverage. The code layer inspected 56 files (TypeScript types, component implementations, Storybook stories) alongside the visual inventory from production screens.
- 14 of 32 components built (44% coverage), 13 of those 14 fully feature-complete
- 24 of 32 components designed in Figma, leaving 8 with no design at all
- 10 components designed but never built, including Card (95+ instances), Select (95+), Menu (200+), Tabs (38+)
- 1 critical component missing entirely Table, blocking 720-1,200 nested instances across 32-42% of screens
Screenshot audit vs API audit
Screenshot-based auditing found 400-700% more component instances than API-based structural auditing. The API saw 10-15 components per page. Screenshots revealed 50 on average.
The reason: the API counted component instances but missed rendered nested content: chips inside cards, icon buttons inside table rows, badges inside list items.
I ran a manual check against the screenshot results to verify accuracy. The counts held up, but a naming problem emerged: the AI and the design system did not always use the same vocabulary. The AI would identify a "select" where our kit called it a "dropdown". Counts were right, labels drifted. This is where a project-level CLAUDE.md becomes essential. A shared glossary that maps the AI's generic terms to the team's component names, so the audit output matches the language designers and engineers actually use. Even with that, some bending remains; our kit uses "dropdown" but the broader industry says "select", and the AI splits the difference.
4-7x more components discovered via visual inspection
Table: the highest-impact gap
Originally estimated at 4 instances from the first 31 screens. A recount across all 57 screens revealed 18-24 instances with 6 distinct patterns: status matrices, data tables, planner grids, calendar grids, comparison grids, forecast tables.
Each table contained 40-50+ nested components: sortable headers, row selection, custom cell rendering with chips, badges, icons, editable fields. No design system alternative existed. Teams used legacy components or built custom implementations. This single component was the highest-impact gap in the entire system.
| Component ↓ | Instances | Screens | Status | |
|---|---|---|---|---|
| Icon Button | 270+ | 91% | Built | |
| Button | 82+ | 95% | Built | |
| Menu | 200+ | 82% | Designed | |
| Select | 96+ | 82% | Designed | |
| Card | 95+ | 68% | Designed | |
| Table | 18-24 | 32-42% | Missing | |
| Navigation Bar | 22 | 100% | Missing |
Deep dive: Card, 12 patterns, no shared contract
Card was the component I chose for the full pattern analysis. The audit analysed 95 instances across all product types. High enough usage to stress-test every phase of the methodology and surface real structural differences.
From those 95 instances, a cross-pattern matrix narrowed the field to 12 distinct patterns, built mainly from title, meta, chips, avatar person, and avatar apps. These 12 were the most flexible for every product to reuse. Running them through the matrices revealed that none shared a fixed layout contract. No two patterns had the same content slot shape.
From audit to tool
Halfway through the audit I noticed I was doing the same thing over and over. Same questions, same matrices, same severity calls. The methodology was disciplined enough that another designer could follow it. So I asked a different question: why not let an AI follow it?
I rebuilt the methodology as a Claude Code slash command. A single Markdown file that drops into a project's .claude/commands/ folder. A teammate adds screenshots, types /design-audit, and the same seven-phase audit runs with the same severity thresholds and report structure.
Two attempts
First version: a multi-file plugin with separate reference files, a templates folder, an examples folder. Technically cleaner. Installation friction was high. Teammates needed to register a marketplace, install the plugin, learn a separate convention.
Second version: a single Markdown file matching the team's existing slash command conventions. No frontmatter. ## Input with $ARGUMENTS, ## Task, ## Rules, ## Section Order, ## Output Structure, ## Success Criteria. British English, no em dashes. Drops in next to the existing commands. Zero new conventions to learn.
The single-file version embeds the full rubric in one place a designer can read and adjust without navigating a folder tree. Adoption was immediate.
Iterating with real use
First production run caught most of the patterns it was supposed to. It also missed one I cared about: a page header pattern (back button, divider, icon, title) that appeared on dozens of screens but was never flagged.
Root cause: the command counted "1 Icon Button, 1 Divider, 1 Icon, 1 text" and moved on. It never noticed those four elements formed a recurring composition. The trigger threshold was built for components, not compositions.
The fix was a new Phase 0, Composition Pattern Scan, that runs before component inventory and explicitly looks for 16 common page-level compositions: page header, section header, filter bar, action bar, toolbar, form footer, empty state, loading state, error state, modal header, modal footer, list item, sidebar item, card header, tab bar with actions, breadcrumb header.
Every miss becomes one more line in the checklist. The command is a living document. Each audit makes the next one sharper.
The prompt
This is the /design-audit command in full. Copy it into a .claude/commands/design-audit.md file in your own repo, adapt the design-system names to yours, and run it. It is meant to be read, edited, and added to by the whole team.
# Design System Audit
Audit UI screenshots against the EDS design system and produce a structured Markdown report covering component patterns, variant gaps, and design recommendations.
## Input
$ARGUMENTS
## Task
Take the screenshots attached above (and any extra context provided) and run a complete design audit following the EDS 2.0 methodology. Produce a single Markdown report saved to the location below.
If fewer than 3 screenshots are provided, run a screen-level audit (lighter, focused on the visible patterns). If 3+ screenshots are provided, run a multi-screen audit with inventory, pattern analysis, and variant gap detection.
## File Location
Audit reports: `apps/design-system-docs/audits/{YYYY-MM-DD}-{scope}.md`
Example: `apps/design-system-docs/audits/2026-05-22-card-deep-audit.md`
If the `audits/` folder does not exist, create it. The `{scope}` segment should be short and descriptive: `card-deep-audit`, `acquire-app-audit`, `multi-app-overview`, etc.
## Rules
- **Only describe what is visible in the screenshots** - do not invent components, patterns, or instance counts.
- **Always include ASCII anatomy diagrams** for each identified pattern - they are the most important deliverable per pattern.
- **Count instances exactly** - report what is visible. Append `+` only when more instances are likely on unanalysed screens.
- **Use British English** throughout (colour, behaviour, summarise, centre, organisation).
- **Never use em dashes**. Use hyphens with spaces ( - ) or en dashes instead.
- **Tag every finding with severity** (P0 / P1 / P2 / P3) - never leave a finding untagged.
- **Recommend existing primitives over new components** when patterns overlap (e.g. an "alert card" pattern should use Banner, not Card).
- **Never invent severity** - apply the rubric below; do not soften or harden findings based on tone.
## Severity Rubric
Every finding, gap and recommendation carries a severity tag.
- **P0 Critical**: Blocks production, no workaround. Triggers: component missing AND used in 50+ instances, OR variant gap forcing every consuming app to build custom, OR accessibility violation.
- **P1 High**: Major gap, inconsistent workarounds across apps. Triggers: 10-50 instances OR gap appears across 3+ applications OR pattern in 30%+ of screens with inconsistent implementations.
- **P2 Medium**: Affects multiple screens, acceptable workarounds exist. Triggers: 3-10 instances OR appears in 1-2 applications.
- **P3 Low**: Edge cases, single-app patterns, easy workarounds. Triggers: 1-2 instances OR single application OR genuinely app-specific use case.
Severity emojis: P0 red circle, P1 yellow circle, P2 green circle, P3 white circle.
## Component Categories
Use these eight categories when inventorying. Always assign each observed component to one category.
- **Form Inputs**: Text Input, Text Field, Text Area, Search, Select, Autocomplete, Combo Box, Checkbox, Radio, Switch, Slider, Date Picker
- **Buttons and Actions**: Button, Icon Button, Button Group, Split Button, FAB
- **Data Display**: Table, Data Grid, List, List Item, Card, Chip, Badge, Avatar, Tooltip
- **Navigation**: Navigation Bar, Sidebar, Breadcrumb, Tabs, Link, Menu, Pagination, Stepper
- **Feedback and Messaging**: Banner, Snackbar, Dialog, Alert, Popover, Progress Bar, Spinner, Skeleton, Empty State
- **Layout and Structure**: Divider, Accordion, Container, Grid, Stack, Spacer, Section
- **Date and Time**: Date Picker, Time Picker, Date Range Picker
- **Surfaces**: Side Panel, Drawer, Bottom Sheet, Popover, Sheet
Use the canonical name from this list, not the consumer's casual name (e.g. "Dropdown" gets recorded as **Select**).
## Usage Frequency Buckets
When reporting screen coverage, use these buckets.
- **Very High**: 75%+ of screens
- **High**: 50 to 75%
- **Medium**: 20 to 50%
- **Low**: 5 to 20%
- **Very Low**: under 5%
## Composition Patterns to Always Check
Page-level compositions are recurring layouts made *of* components - they are not components themselves, so the inventory phase will miss them unless you check for them explicitly. Always scan every screen for the following compositions, regardless of how few times they appear. Each one found counts as a pattern in its own right.
- **Page header**: back button, optional divider, optional icon, title, optional trailing actions. Pattern: `back | [icon] Title ... [actions]`
- **Section header**: title, optional subtitle, optional trailing actions or count.
- **Filter bar**: search input, filter chips, sort control, view toggle.
- **Action bar**: contextual actions that appear on selection (e.g. "3 selected | Edit | Delete | Move").
- **Toolbar**: grouped icon buttons, usually with dividers between groups.
- **Form footer**: divider above, cancel button (left or ghost), primary action (right).
- **Empty state**: icon or illustration, heading, supporting message, optional CTA.
- **Loading state**: spinner or skeleton, optional message, optional progress.
- **Error state**: icon, heading, supporting message, optional retry CTA.
- **Modal header**: title (left), close X (right), optional subtitle below.
- **Modal footer**: divider above, secondary action, primary action.
- **List item**: leading element (avatar, icon, checkbox), main content (title + supporting text), trailing meta or action.
- **Sidebar item**: icon, label, optional count or badge, optional active state.
- **Card header within Card**: title, optional metadata, optional trailing menu or actions.
- **Tab bar with trailing actions**: tabs (left), action buttons (right of the tab strip).
- **Breadcrumb header**: breadcrumb path, optional title below.
For each composition found, document it the same way you would document a component pattern: name, app, count, ASCII anatomy, slots used, severity if there are inconsistencies across screens. Add them to a dedicated **Composition Patterns** section in the report, separate from the component pattern analysis.
If a composition pattern repeats across 3+ screens with material differences, flag each variation as a separate composition variant and tag with severity per the standard rubric.
## Audit Phases
Follow these phases in order. Do not skip phases even if a screen looks simple.
### Phase 0 - Composition Pattern Scan
Before counting components, scan every screen for the compositions listed above. Document each one found using the same anatomy format as component patterns. This phase is mandatory even if only one screen is provided.
### Phase 1 - Inventory
List every visible UI element across all screenshots. Group by component type. Count per screen and across the whole set. Output: a count table with component, total instances, screens using, and coverage percentage.
### Phase 2 - Pattern Identification
For each component appearing 5+ times OR on 20%+ of screens, identify the distinct patterns it takes. Two instances are the same pattern only if they share layout, content slots, and interaction behaviour.
For every pattern, document: pattern name and host application; instance count and layout; ASCII anatomy diagram (always include); properties (border, background, radius, shadow, click behaviour, hover state); CSS characteristics (best-guess from screenshot); content slots used.
Sub-pattern threshold: only declare a new pattern when layout, interaction, or semantic visual treatment differs. Pure content differences are NOT a new pattern.
### Phase 3 - Cross-Pattern Analysis
For each multi-pattern component, produce three matrices: Visual Properties, Interactivity, and Content Slots.
### Phase 4 - Variant Gap Analysis
For each component: list variants observed in production; list variants documented in the design system; compute the diff (Observed minus Documented equals Gaps); score each gap with the severity rubric.
### Phase 5 - Component Recommendation
For each multi-pattern component, write a recommendation section. Always include a "What NOT to build into [Component]" section listing features that should be left to composition.
### Phase 6 - Severity and Prioritisation
Tag every finding from phases 2 to 5 with P0 / P1 / P2 / P3.
### Phase 7 - Assemble Report
Use the Output Structure template.
## Success Criteria
- The report follows the section order exactly.
- Phase 0 was performed before component inventory, and every composition found is documented with an ASCII anatomy diagram.
- Every pattern includes an ASCII anatomy diagram.
- Every finding is tagged with P0, P1, P2 or P3 using the rubric.
- Instance counts and screen coverage percentages cite what is visible - no fabricated numbers.
- British English used throughout.
- No em dashes anywhere in the output.
- Component recommendations take a position with justification, not hedged language.
- Patterns overlapping with existing EDS primitives (Banner, Chip, etc.) are redirected to the primitive, not given new variants.
Impact
- 44% to 72% coverage path Implementing the 10 designed-but-not-built components would nearly double design system coverage
- 4 P0 and 7 P1 components prioritised with severity-tagged evidence, replacing reactive roadmap decisions
- 720-1,200 blocked instances surfaced by identifying Table as the single highest-impact missing component
- 56 files, 122+ props, 80+ stories audited in the code inspection layer
The /design-audit command lets any designer on the team run the same audit on new screens without me. Time-to-audit went from days to minutes per screen set. The methodology is now reproducible across teams, projects, and future hires.
Composition patterns surfaced design system gaps that were invisible to per-component audits. The screenshot-first methodology was validated as 4-7x more accurate than API-based structural auditing.
What I learned
Skills are a team craft
Building the audit skill taught me how to make AI tooling for a design system team. The goal was never a personal shortcut. It was a shared capability any designer could run, review, and trust. That means writing the skill in the team's language, matching existing conventions, and keeping the output legible to someone who did not write the prompt.
AI is a living thing
The skill gets better with every contribution and every new model. But it also needs maintenance. Naming drifts, thresholds shift, new patterns appear. We will have to keep updating the skills without losing sight of why we use them in the first place: to make repeatable work fast and consistent, not to replace judgement.
Small teams need leverage
We are one PO, one lead designer, one lead developer, one mobile developer. Between audit, component priority, UI work, testing, and community, the surface area is large. AI gives a small team like ours the ability to get results fast on the repetitive, high-volume parts of the work. But fast does not mean unchecked. Every output still needs a human pass. The speed gain only holds if we keep the review loop tight.