Operations Platform · Case study

Inbound SSO Administration

Everyone thought we were building a form. We were building the thing people reach for when login breaks.

Role

Senior UX Designer, design lead for discovery

Timeline

1 month discovery

Counted

Research

4 moderated sessions, one per role

Counted

Status

[[Shipped / In build / Recommendation adopted]]

Counted = from my work. Observed = from my research sessions. Benchmark = external research, context only. Target = a goal I am measuring against.

A tool nobody wanted to need

Fidelity is retiring an older authentication server. Before it can go, the people who connect applications to single sign on need a tool they can run themselves.

Today, a lot of that work routes through the central security team, by hand, across several older systems. That works until it doesn’t.

The first requirements asked for three things: create a connection, edit it, view it. Reasonable. Also slightly off.

Counted

Nobody started with “create”

What the requirements assumed

Create

Edit

View

Configuration management project.

What the sessions showed

Search

View

Investigate

Validate

Modify

Operational administration platform.

I sat down with four people who touch this system every week. Nobody opened with creating anything. They opened with finding. Which connection is this? Is the certificate okay? What breaks if I change it?

Observed

[[X of 4]] started their walkthrough with search. Creating a connection is rare. Finding, checking, and fixing is the job.

3 verbs became 5

Counted

[[X of 4 started with search]]

Observed

Frequency data is from participants, not system logs. See the impact section for how I am closing that gap.

Four people, four links in the chain

Observed

1 Reviewed requirements and architecture documents.

Counted

2 Mapped systems, owners, and dependencies.

Counted

3 Ran four sessions.

Counted

4 Synthesized themes and gaps.

Counted

5 Recommended MVP boundaries and risks.

Counted

Deployment

onboarding, configuration, validation

Deployment leadership

risk, governance, efficiency

Testing and support

troubleshooting, validation, supportability

Production support

investigations, operational visibility

I asked each person to walk me through their last real change and their last real incident, not a hypothetical. [[Confirm method]] I played the synthesized requirements back to them to check I had it right. [[Confirm playback step]]

How much can four people carry?

Four people is not a statistic. It is one voice for each link in the chain. Nielsen and Landauer’s model says each person in a similar group finds about 31% of the problems, which puts four people around 77%. My four were four different roles, so I treat that as a ceiling, not a promise. It is strong for direction and for spotting what is missing. It is weak for counting how often things happen. So I treated findings as strong signals and went back for numbers where a decision depended on them.

Benchmark

Who owns what, finally on one page

Deployment

Deployment

owns onboarding, configuration, and validation. Outputs: configuration records, Entity ID validation.

Testing

Testing

validates configurations across environments before production. Outputs: test results, environment validation.

Support

Support

investigates production issues and certificate health. Outputs: investigation reports, incident resolution.

Architecture

Architecture

defines system boundaries, data models, and integration contracts. Outputs: architecture decisions, API contracts.

For the first time, product, architecture, deployment, and support were looking at the same map.

Five things I would stake the roadmap on

Counted

1. Search is the main job. Design for search first and creation second.

[[High / Medium]]

Observed

2. Entity ID runs through everything. It stays visible in every workflow.

[[High / Medium]]

Observed

3. Visibility beats editing. People needed to see more before they needed to change more.

[[High / Medium]]

Observed

4. Support needs more than the plan allowed. It earns investment beyond the original scope.

[[High / Medium]]

Observed

5. One Entity ID can serve many products. The data model has to handle one to many.

[[High / Medium]]

Observed

One thing, three names

Observed

Product model

Connection

how the system stores it.

User model

Client, vendor, support request

how people talk about their work.

System model

Entity ID

the identifier everything hangs on.

Same object, three names. If the interface picks one, two groups of people get lost.

Design direction: search that lets people start from whichever name they know, and shows the other two right away. [[Confirm this matches the concept]]

Four gaps that would have shipped

Counted

Gap

What I heard

Design response

Risk

Gap

Gap

Delete connection

What I heard

What I heard

Teams have disabled production configurations by accident [[number if known]]

Design response

Design response

Protected deletion with a dependency check

Risk

Risk

High

Gap

Gap

Delete before defederate

What I heard

What I heard

Deleting first can orphan user identities

Design response

Design response

Dependency validation before any destructive action

Risk

Risk

High

Gap

Gap

Certificate visibility

What I heard

What I heard

Connection status and certificate status are separate states

Design response

Design response

Certificate health with its own indicator, never color alone

Risk

Risk

Medium

Gap

Gap

Connection search

What I heard

What I heard

Search is the primary workflow

Design response

Design response

Search first with progressive disclosure

Risk

Risk

Medium

4 gaps found, 2 high risk

Counted

1 classification became 3 dimensions (ID type, SSO type, SAML profile)

Counted

What I recommended, and why in that order

MVP 1, Migration Readiness.

Counted

Goal: deployment teams administer connections on their own.

Includes: search, view, create, modify, delete, reactivate, entity validation, audit history, connection metadata, multi product support.

Business outcome: less dependence on the central security team for routine changes.

User outcome: onboard and manage connections without escalation.

Risk reduction: removes the migration blocker, because the old server cannot retire until this exists.

MVP 2, Operational Confidence.

Counted

Goal: make production issues faster to understand and fix.

Includes: certificate health, advanced search, reporting and dashboards, performance insights, advanced audit views.

Business outcome: lower mean time to resolution for production SSO incidents.

User outcome: support can investigate without deep platform expertise.

Risk reduction: catches certificate and health problems before users do.

MVP 1 exists because of the retirement date. MVP 2 exists because a tool you cannot diagnose in is a tool support will not trust.

Counted

What changed, and what I will measure

What discovery changed

4 gaps found before build.

Counted

2 high risk gaps caught.

Counted

Workflow verbs 3 to 5.

Counted

Classification 1 to 3.

Counted

[[Number of original requirements changed or added]].

Counted

Measurement plan

Metric

Baseline

Target

Evidence

Metric

Metric

Time to find a specific connection

Baseline

Baseline

[[minutes, self reported, n=4]]

Target

Target

At least 2x faster

Evidence

Evidence

Target

Metric

Metric

Routine requests escalated to the central security team per week

Baseline

Baseline

[[number]]

Target

Target

Routine escalations eliminated

Evidence

Evidence

Target

Metric

Metric

Accidental production disablements in the last 12 months

Baseline

Baseline

[[number]]

Target

Target

Zero

Evidence

Evidence

Target

Metric

Metric

Time to confirm certificate health

Baseline

Baseline

[[minutes]]

Target

Target

At least 2x faster

Evidence

Evidence

Target

Metric

Metric

Hard deadline: server retirement

Baseline

Baseline

[[date]]

Target

Target

Administration live before it

Evidence

Evidence

Target

Why this matters beyond Fidelity

Context, not my results

Uptime Institute (2025): nearly 40% of organizations had a major outage caused by human error over three years, and 85% of those came from staff not following procedures or flawed procedures. This is data center research, and I use it as an analogy for protected deletion.

Benchmark

Keyfactor and Ponemon (2019, 596 US practitioners, vendor sponsored, self reported): 74% reported downtime caused by digital certificates, and 71% said they do not know how many certificates they have. I use it as context for certificate visibility.

Benchmark

An admin tool deserves the same standards

Design requirements

Search is fully keyboard operable with a clear focus order.

Status is never shown by color alone.

Destructive actions are reversible, checked, or confirmed, following WCAG 2.1 success criterion 3.3.4.

Benchmark

Search result counts are announced to screen readers.

Labels use plain language, and every field says what it is in the words the team actually uses.

What other teams can borrow

the Entity ID relationship pattern

the protected destructive action pattern

independent status indicators

[[Confirm whether these are in the shared pattern library]]

The moment it turned

[[One paragraph in my voice: the moment Product’s view changed, who pushed back, and what moved them.]]

What I would do differently

[[One short paragraph.]]

Do you have any project ideas you'd want to discuss?

Do you have any project ideas you'd want to discuss?

Do you have any project ideas you'd want to discuss?