About
Artist first, product designer second.
I design AI products companies use.
Six years turning fuzzy ideas into shipped AI product design — from first thought to product–market fit —
at Apple, and now with startups as an independent designer.
Before product design there was photography. Chasing light across the world.
It taught me the two things AI can’t fake: taste and intuition.
Everything on this site is the same move repeated: something large and tangled arrives, and leaves as something a person understands in a second.
Every project runs the same arc: brief, concept, functional prototype, developer-ready design.
I start with research — interviews, shadowing real users, learning the tool as its newest user.
Then I map journeys and flows before any screens, and iterate through usability feedback to high fidelity.
I prototype in working code so stakeholders click instead of imagine.
And I hand engineering a style guide, not a pile of mockups.
Most of what you’ve seen here, I designed alone.
When a project needs more hands, I bring people I’ve worked with and trust — and I stay accountable for the result.
One point of contact, one standard.
“Having those previously missing forms restored will be of major impact to the operations teams.”
— Product owner, Apple, on release
A note on confidentiality
The Apple work on this site is internal and unpublishable, so there are no screenshots,
and internal product names have been replaced.
The thinking is mine to show; the products are not.
Apple — Station 01 / 05 — Define
Form Builder
An internal Apple platform where managers build evaluation forms and score customer interactions — 160,000 users, five million evaluations a year.
Apple · 2022–2023 · Only designer
When I joined, I was the newest person using a tool that 160,000 people had already learned to work around.
The platform was old. Outdated, confusing, hard to navigate.
Years of complexity had accumulated without a meaningful redesign.
And yet people used it every day.
That was the interesting part.
The product was clearly broken. But the organization had adapted around it.
The people who had used it for years had learned its quirks.
- They had memorised the paths.
- They knew the vocabulary.
- They kept cheat sheets.
I was the newest user, so I had none of those workarounds.
And as the only designer, I paid very close attention to every moment where I thought:
The first thing I noticed was what happened when you opened the app.
You landed directly on a list of forms.
- No dashboard.
- No recent work.
- No sense of where you’d left off.
You had to already know which form you were looking for.
For someone who had used it for years, that was manageable.
For me it was immediately confusing. And that difference became the starting point.
Two experiences were hiding inside one product.
Veterans weren’t using the interface anymore — they were following memorised paths. New users had none of that.
I wasn’t just learning the product. I was learning what 160,000 people had been forced to learn.
How do you design something new while still making it feel familiar?
I couldn’t throw the old experience away and build what I thought a modern product should look like.
That would have solved the usability problems and created a worse one.
160,000 existing users would have to relearn everything.
So the question stopped being what to change, and became:
How much can I change before something familiar becomes something people have to relearn?
That became the principle behind almost every decision I made.
I started with the first few seconds.
The system already knew what you’d worked on, what you’d left unfinished, which forms you used. None of it was visible.
So I brought that state forward.
I wasn’t adding a capability — I was surfacing what the system already had.
Make the product feel like it knows why you’re here.
Then the bigger problem. I could have changed the workflow, but experienced users had years of muscle memory in it.
So I made a deliberate decision.
Change the interface. Keep the sentence.
- The vocabulary stayed.
- The order of the steps stayed.
- The surrounding experience changed.
“This looks different… but I know what to do.”
The goal was never to make the product beautiful for me.
It was to make it better without making 160,000 people start over.
The platform needed consistency. But people needed freedom.
Questions lived in a shared library, which kept evaluations comparable across teams.
But managers needed to create new ones halfway through building a form.
Control the library and creation becomes painful.
Open it completely and the library fragments.
So I changed the relationship: questions and answers became independent objects.
A question created in the moment joined the library immediately, and could be reused anywhere.
Creation stayed free — and instead of fragmenting the library, it grew it.
Then I had to make the complexity understandable.
Categories, questions, answers, scoring, conditional logic, metadata. It was becoming a small programming environment.
Treat the form like an outline.
A collapsible tree on the left for the structure. A panel on the right for the details.
- The tree answered: “What does this form ask?”
- The tabs answered: “How does it behave?”
And then AI entered the conversation.
The platform was built for humans evaluating humans. But an AI evaluator was coming.
Take a question like “Did the employee reassure the customer?”
A human evaluator understands that. A model needs more.
- What actually counts as reassurance?
- What does a pass look like?
- What does a fail sound like?
So I designed an AI layer into the evaluation itself: a plain-language rule for each answer, examples of passing, examples of failing.
A manager could test it against a real interaction and see the model’s reasoning.
The important part wasn’t the AI. It was who was teaching it.
The people who understood the evaluation were the managers.
They shouldn’t need to become ML engineers to teach a model what good looked like.
Looking back, the biggest change wasn’t the UI. It was that the product finally acknowledged what it already knew.
But I also learned when to stop.
I’d explored form templates and richer tagging. Both reasonable. Nobody needed them badly enough.
The product already asked a lot from new users.
Anything that didn’t earn its place was one more thing to learn. So I cut them.
Sometimes making a product better means removing things, not adding them.
When you redesign a decade-old product, you’re not designing for a blank canvas.
You’re designing around everything people have already learned.
So I stopped asking “How would I design this today?”
And started asking:
“What can I change without making people forget everything they already know?”
New enough to be better. Familiar enough to be trusted.
“Having those previously missing forms restored will be of major impact to the operations teams.”
— Product owner, on release
Apple — Station 02 / 05 — Evaluate
Evaluator
The internal Apple tool where managers search, replay and score customer calls and chats, with a team’s whole day laid out as a timeline.
Apple · 2020–2021 · Second designer, supporting
This was my first project at Apple. I was the second designer on it, and the newest person in the room.
The tool is where managers actually do the evaluating.
Everything I designed afterwards either feeds this screen or reads from it.
A manager searches every customer interaction — any channel, any filter — finds one conversation, listens or reads it,
and drops markers on the moments that matter.
- Rudeness at 9:29.
- A missed transfer.
- A perfect save.
The screen also holds the whole day: a team’s schedule as parallel tracks — calls, holds, transfers, breaks, training, offline.
So a conversation is never judged out of context.
That mattered more than I expected.
A difficult call at the end of a six-hour queue is not the same call at nine in the morning.
When a manager evaluates a conversation, the evaluation form appears beside the transcript.
Built in the platform from the first study. Rendered live in this one.
- A moment becomes a marker.
- A marker becomes evidence.
- Evidence becomes a score.
My role here was supporting — specific requirements, specific screens, helping the lead designer.
I include it because it’s the room the rest of the system lives in.
And because starting here, as the newest user of everything, is what taught me the loop.
Apple — Station 03 / 05 — Understand
Tracker & Coaching Product
A performance platform for Apple managers: one page per team member, designed from a blank page to replace two legacy tools. Used worldwide.
Apple · 2023–2024 · Lead designer
This one I started from a blank page, as the lead designer.
A performance platform for managers, replacing two legacy tools.
The first thing I noticed was that the data already existed.
The data existed. The picture didn’t.
One tool held the team’s metrics. The other held individual surveys and feedback.
Both showed data about a person, in different places, in different shapes.
And the manager assembled the picture in their head. Every week. For every team member.
That’s slow. But slowness wasn’t the real cost.
You can’t see a pattern in your head.
A dip in quality that lines up with a run of bad surveys that lines up with a change in schedule.
That’s the thing a manager most needs to notice — and precisely what two separate tools make invisible.
So I organised the platform around the person.
Pick someone on your team and everything about them is on one page, under four lenses:
- Performance.
- Evaluations.
- Surveys.
- Peer feedback.
The design rule for that page was blunt:
A manager should know what’s going on with this person within a second of landing.
Not after scrolling. Not after opening a tab.
Which meant deciding, for every piece of data, whether it earned the first view or lived one click down.
Then: surface patterns, not points.
Every headline number is shown twice at once — against the person’s own history, and against their group.
An adoption rate of 50% means something different if the team is at 70% than if it’s at 45%.
And beneath the score sits the structure that produced it: the criteria, each with its share of the gap.
So when the pattern says something is changing, the tree says where.
The conversation moves from “you’re at 51%” to “resolution handling is the thing, and it started three weeks ago.”
Then the most human data on the page, and the least reliable.
A colleague flags a case. Useful — but not automatically fair.
So peer feedback doesn’t land on a record raw. The manager reviews it, marks it valid or invalid, assigns a root cause.
Only then does it become part of the picture.
Feedback stays fast to give and slow to count.
“Great job collaborating, listening to everyone’s thoughts, and turning them into reality.”
— Product manager
Apple — Station 05 / 05 — The loop, applied to AI
Chatbot Evaluator
Metrics for Apple’s customer-facing AI chatbots: defining what a failed conversation is, then measuring 1.8 million chats in sixty days.
Apple · 2024–2025 · Lead of the metrics portion
By this point I’d spent years designing how humans evaluate humans.
Then the same question arrived about a machine.
What does a failed AI conversation look like?
I owned the metrics side — the dashboard and every screen that measures how the bot is doing, and decides what it learns next.
With a human employee, you know what a bad call is. Someone listens, scores it, writes down why.
A bot handles 1.8 million chats in sixty days.
Nobody is going to listen.
And “the bot did badly” isn’t a thing a system can count. It’s a judgment.
So before a single chart, the team needed a definition of failure a machine could detect.
I worked with the AI engineer to turn “unsuccessful” into things the system already observes.
- The customer abandoned the chat.
- The chat timed out.
- The bot handed off to a human.
- Or the customer said it didn’t help.
Failure is a set of events, not an opinion.
Then I ran into something I didn’t expect.
We couldn’t get the whole team to a definition everyone trusted.
Every signal had a counter-argument. A hand-off to a human can be a failure, or exactly the right call.
So we cut back to what was undeniably true — total chats, and chats by topic — and held the rest for a later release.
I’d rather ship a small number people trust than a big one they argue with.
The dashboard’s credibility was the product.
Then the design problem itself. A dashboard that says “60K failed” is a mood.
The job was to make every number a doorway.
Total unsuccessful, then by topic, then by subtopic, then the individual conversations — each with the customer’s own reason attached.
Three clicks from a company-wide figure to one person saying “you redirected me to a page I’d already read.”
One more thing the bot needed that a person doesn’t.
It changes constantly — every retrain is a new release. So every metric is filtered by version.
Pick a release, see where failures rose or fell against the previous one,
and judge a retrain by evidence rather than impression.
Looking back, this project is the mirror of everything before it.
- Define what good looks like.
- Evaluate.
- Understand.
- Act.
Only here the employee is a model, the evaluators are events, and “act” means retrain.
Having designed both sides of that mirror is the through-line of everything on this site.
Fortify AI — 2026 — Lead designer
AI Security Platform
A zero-trust security and governance platform for AI agents. I led design end to end — research, architecture, visual system, prototypes, style guide.
Fortify AI · 2026 · Lead designer · No product imagery
This is the first one that wasn’t Apple.
Fortify registers the AI agents an organisation runs, enforces policies on what they may do,
surfaces incidents, and produces audit evidence.
The interface I inherited had eleven top-level sections. Every one of them was real.
Together they described the engineering, not the work.
A security analyst opening the product at seven in the morning has three questions:
- Is anything wrong?
- What is it?
- What do I do?
None of the eleven sections was an answer to any of them.
The dashboard greeted you by name and then showed five equal cards.
Which is a way of saying nothing is more important than anything else.
So I mapped every section to one of three jobs — register and control, monitor, audit — and collapsed the navigation to five.
And the dashboard got a hero: one security posture score, with everything else deliberately smaller.
One glance, one decision.
Then came the part I got wrong first.
Every serious competitor in AI security ships dark. Navy or charcoal, one electric accent, geometric sans.
I audited a dozen of them and nine benchmark developer products. The pattern was unanimous.
So the second direction went dark. Off-black surfaces, a signal-lime accent nobody had claimed, an editorial serif for the hero number.
It looked like it belonged in the category.
After a month living in it, it read as a marketing site, not an operator’s tool.
The dark glass flattened hierarchy, the accent fought the severity colours, and the whole thing was louder than the data.
So the third direction reversed it. Pure white surfaces, a cool blue-grey sidebar,
zero radius on everything, so the interface reads as a terminal rather than a landing page.
Two typefaces with strict roles: a humanist sans for anything a person wrote, a monospace for anything the system knows.
Severity in earth tones, so a single red pill keeps its weight on a page full of data.
The aesthetic isn’t softer. It’s quieter.
Going light in a dark category was the riskiest call in the project.
It was also the one that made the product look like it was built by people who use it.
I started in Figma and moved to building prototypes directly with AI — working HTML the founders could click through.
Three full visual directions in four months would not have been possible in static mockups.
I used an AI competitive audit as a starting brief. Its recommendation was dark and lime. The product ended up light.
Apple — Station 04 / 05 — Act
Coaching Product
The tool that turns a flagged moment into a coaching plan — and the only station in the system where employees see themselves.
Apple · 2025 · Supporting designer
This is where the loop finally closes.
Everything before it measures. This one does something about it.
A manager creates a plan for a team member — tasks, milestones, a journey the person can see and follow.
A moment flagged in a call can become the seed of it, with the exact timestamps attached.
So coaching starts from evidence instead of memory.
There’s one other thing that makes this station different from all the others.
It’s the only one where employees see themselves.
They have a profile. They can follow their own progress.
Every other tool in the system looks at people. This one is looked at by them.
I joined an existing design and pushed it forward — new features, clearer requirements,
and a move from static mockups to shareable prototypes built with AI.
Stakeholders clicked instead of imagining. That changed the conversations.
There was a smaller chapter alongside it: a light redesign of an internal tool for a different audience,
a clean refresh to bring a dated interface up to the present.
Sometimes the job is not to rethink. It’s to make something feel cared for again.