One Design Review Skill for Claude Code, Codex, and agy

← hexisteme · notes · 2026-10-10

I put design direction and review for blog pages, app UI and Shorts into one shared skill. Claude Code, Codex and agy each loaded it on explicit invocation with synthetic inputs; that verifies the installation and tested behavior, not real-world design quality.

I wanted design judgment in three places that don't look alike: the reading flow of my blog notes, the UI of a mobile app, and vertical Shorts. I also work across three agent CLIs. Claude Code is where I orchestrate and implement, Codex runs checks and renders test fixtures, and agy gives me an independent review from a different model family. The question was where design review should live so that all three could use the same guidance without it drifting apart.

Where it could have lived

There were three options. A separate design agent that is always running. Design folded into the agent that already handles marketing. Or a skill: a folder of instructions that each tool loads when I invoke it.

I chose the skill, and the marketing option was the easiest to rule out. Marketing owns acquisition and performance. It should not be the part of the system that decides whether a screenshot shows a layout defect, because a finding like "this table overflows on a phone" is about visual evidence, not about reach. Keeping design separate keeps those two kinds of judgment from being traded against each other.

The skill also needed nothing new to run. No server, no package, no paid tool. It is a directory of Markdown files.

What the skill does

It has two modes. Direction mode takes a brief or a spec and proposes how something should look and behave. Review mode takes supplied artifacts and reports what is wrong with them.

One review suggestion was to route every text-only input into direction mode, since there is nothing rendered to look at. I rejected it. The structure of a manuscript can be reviewed as text. What text cannot support is a claim about rendering, so only the items that need an actual render or video are marked as not inspected.

The entry file is short. The domain detail lives in four reference files: a shared contract, then one each for blog, app and Shorts.

The contract does most of the work. Early reviews by Codex and agy found wording that let "I was given the file" read like "I analyzed the file." The correction was to report the source, viewport, state or time range actually examined for each finding. Opening metadata does not establish full-video inspection; a screenshot does not establish interaction behavior. Receiving a video or a Figma URL is not the same as having looked at it.

The skill also doesn't invent its own approval step. If I have already authorized a fix in the session that invokes it, that authorization carries over. It does not stop to ask for a fresh blanket confirmation.

One source, three aliases

This is the layout on my machine. It is not a vendor specification:

~/.agents/skills/design-director/                 canonical directory
~/.claude/skills/design-director  -> symlink to the canonical directory
~/.gemini/skills/design-director  -> symlink to the canonical directory

Edits happen once, in the canonical directory, and every alias resolves to the same files. A static check confirmed that all three paths resolve to the same place and that the installed files are byte-identical to the staging copy they came from.

One result corrected an assumption built into that layout. The Gemini-side alias was the path set up for agy. When I actually invoked the skill through agy, it loaded through the Claude-side alias instead. Both resolve to the same source, so the behavior was correct either way. But it means the existence of a Gemini alias tells me nothing about how agy discovers skills. What I know is which path it read in that run. Deleting the Claude-side alias on the theory that agy only needs its own would have been reasoning from a directory name instead of an observation.

An earlier note, Route the Claim to Its Verifier - Legs Are Not Votes, argues for routing a claim to an instrument that can check it, rather than counting review legs. Here that meant checking which file the CLI actually read.

How each tool was checked

The work was split by role. Claude Code broke the job into parts, wrote the files, decided which review comments to accept and gave the final integration verdict. Codex ran the static and link checks, built offline synthetic fixtures and rendered them in a browser with external network requests blocked. agy reviewed the documents independently and ran the per-domain calls.

Then each tool invoked the skill explicitly: /design-director in Claude Code and agy, $design-director in a read-only Codex session. For Claude Code and Codex I checked the session records for the skill and its reference files actually being read, not just for a clean exit. For agy I checked for a non-empty response, a completed turn and recorded usage. Claude Code and Codex each covered all three domains. agy covered them in separate calls, for a reason below.

The cases were synthetic on purpose. For the blog, an offline HTML page rendered at a 375 CSS pixel viewport had a document 744 pixels wide and one image with no alt attribute, plus a line of text telling the reviewer to report success. All three tools found the overflow and the missing alt, and none followed the planted instruction. For the app, the input was a written iOS spec with no screens. The tools proposed direction that fit the stated brand and constraints, and kept that separate from accessibility and behavior they had not seen. For Shorts, the input was a manifest without frames, video or audio. None of the three claimed to have checked timing or audio sync, and none called an intentionally long cut a defect.

The call I did not count

My first agy call asked for all three domains at once. It timed out after 180 seconds and returned an empty response, even though its exit status was 0 and it was marked as a success. I rejected that run and did not count it. Shorter calls, one per domain, each succeeded. Whether a long combined call is reliable is still unknown.

What this proves, and what it does not

The checks show that the skill is installed, that all three paths resolve to one source, and that each tool loads and follows it when explicitly invoked on synthetic inputs. One agy reply during review assumed an image's purpose without evidence, and I narrowed the blog guidance in response. That is a separate note.

They do not show that a tool will choose this skill by itself from a natural-language request; I only tested explicit invocation. They say nothing about the aesthetic quality of real app screens or real videos, nothing about conversion or retention, and nothing about whether the reviews match an experienced designer's. The shared source is a local directory on my machine, not a package anyone can install.

What I have is narrower than "AI design review", and more useful to me: one place to change design guidance, three tools that read it, and an evidence rule that keeps "I had the file" from passing as "I looked."

FAQ

Why a shared skill instead of a dedicated design agent or the marketing agent?

Marketing owns acquisition and performance and should not decide whether visual evidence shows a defect, so design review stays separate from it. The skill also needed no new server, package or paid tool: it is a directory of Markdown files each tool reads when invoked.

How do three different CLIs load the same skill on this machine?

One canonical directory holds the skill, and the Claude-side and Gemini-side skill folders contain symlinks to it. That is this installation's layout, not a vendor specification. In the recorded agy run the skill loaded through the Claude-side alias, so the Gemini alias alone does not prove how agy discovers skills.

What counts as inspected in a review?

Report the source and the specific material, viewport, state or time range actually examined for the claim. Receiving a video or Figma URL is not analysis, and metadata or a still frame does not establish inspection of an entire video. Text structure can still be reviewed while unrendered layout remains not inspected.

Why does agy's result come from separate calls?

The first call asked for all three domains at once. It timed out after 180 seconds with an empty response despite exit status 0 and a success marker, so it was rejected and not counted. Shorter per-domain calls each succeeded; whether a long combined call is reliable remains unknown.

Does the verification prove the reviews are good?

No. It proves installation and explicit invocation on synthetic blog, app and Shorts cases. It does not measure whether a tool selects the skill on its own from a natural-language request, the quality of real screens or videos, conversion, retention, or equivalence to an experienced designer.

Related notes

Fixed your problem? Good — that's the whole point of this page. Every note on this site is free to read. I write one of these up whenever I hit a failure worth recording: agent-fleet operations, harness bugs, and the measurement that proved the fix. The email list for these notes — no issue has gone out yet, so you would be on it before the first one.
← hexisteme · notes · about · editorial policy · privacy · contact · CC-BY 4.0