Nayya / From SaaS to AI‑nativeNayya · AI‑native
Chapter 01 · Jupiter · July–October 2025

Could chat make benefits legible?

The request was to add an AI chat, but that was a feature request masquerading as product strategy. A chat interface might have made the product feel more “AI,” without addressing users’ long-standing pain points.

My role: product design lead. Co-led strategy with the Director of Marketing, CPO, and PM, with Data and Engineering.
Initial modelChat-first, chat-onlyTurn the product into an “AI” product by making chat the primary interface. First bets: leave planning and benefits.
Discovery inputsSupport, Sales, Marketing, product behaviorPlus executive workshops and review of collateral and drop-off patterns
Learning modelThree playtests per weekPlus a market pilot with about 100 HR professionals and industry stakeholders
Key resultUseful, but not sufficientConversation improved access and explanation; complex answers remained hard to scan and act on
Nayya chat on laptop and phone
01 / Opportunity

The need was not one benefits question. It was a recurring loss of context.

Employees move between employers, plans, life stages, and benefit packages. A new, unfamiliar package turns even an experienced user into a beginner: what does this plan cover, what does that term mean, and what should I do next. A new life stage (a dependent, an illness, retirement) raises the stakes again and brings another round of information and guidance. People rarely become lasting experts in their benefits. They keep needing help as the situation changes.

They also bring priorities that pull in different directions: maximize coverage, keep costs down, or simply see every option before choosing. With that much context and nuance, it was never one question. It was never one answer.

Need 01

Answer

Answer the questions I already have. What does this benefit cover, and what do the terms mean?

Need 02

Relevance

How does this impact me and my situation? How does it apply to where I am right now?

Need 03

Optionality

How else can this be covered? How would it look different in another circumstance?

02 / Product thesis

Generic answers were not enough. The differentiator was personal context.

Users could already search the web for definitions. Jupiter became valuable when an answer combined a plain-language explanation with the user's actual plan details, deductible progress, costs, dependents, or prior recommendation inputs, plus their current situation, past medical history, and upcoming needs.

We started from strong executive conviction

Mainstream AI products had made a blank conversational prompt feel familiar. Leadership believed a similar interface could simplify benefits navigation.

Nayya had never shipped a chat assistant, so the team needed to create the interaction language, answer structure, component system, and behavior model from scratch.

I translated a broad idea into answer types

I mapped recurring needs into bounded patterns such as definitions, cost calculations, coverage explanations, and personalized plan or profile answers.

This created a shared unit of design for Product, Data, and Engineering: not "make chat better," but improve the structure and trustworthiness of a specific answer type.

Anatomy of a personalized answerPrototype · representative
03 / Operating model

Conversational design required a faster learning loop.

Traditional weekly or biweekly sprint feedback was too slow. A response could be technically correct and still fail because of order, length, tone, missing context, or the wrong interaction form.

3×cross-functional playtests every week for roughly two to three months
01Prompt
02Observe
03Adjust
04Replay

I aligned the CTO and CPO on the cadence

Rotating participation brought Engineering, Product, Design, Data, and executives into direct contact with product behavior. It was repetitive, but it surfaced length, tone, and details that only show up occasionally when users ask — issues missed if testing waited until the end.

I specified generated behavior like product UI

I created a text-based AI design specification in Notion with answer formats, templates, guardrails, and golden examples as a concrete target for generated output. We also ingested those golden answers as examples the model could pull from when responding, wired straight into the codebase.

Research & workshops

Research

AI tool market analysis

Workshop

Value chain analysis

Workshop

Feature priority collaboration

Workshop

Tone calibration

Research

Persona & use case analysis

Testing

Play-testing (weekly)

Testing

AI answer optimization

Workshop

Design review and feedback

Working specs that made the answers holdScrollable source documents
Open full document ↗
Jupiter Testing Success Criteria: acceptance rubric for playtest answers
Tone testing feedback scorecard from playtests
Workshop map of AI roles across autonomy and engagement
04 / Validation

Jupiter validated personalized guidance and exposed the limit of chat.

We took that system into the market and tested it with about 100 people: HR professionals, CHROs, C-suite and other company stakeholders, plus brokers, carriers, PEOs, and consumer users.

What users valued

General explanation plus personal context

People responded to answers that defined a term and immediately connected it to their own plan, deductible, progress, and costs.

What clients valued

Answers grounded in their documents

Clients saw value in employees asking about employer-specific materials without leaving the product for generic search.

What didn’t work

Walls of text, missing predictability, and trust beyond tone

Long answers became walls of text users skimmed. Forms had to return for leave and other multi-step work so scope stayed parseable and predictable. Trust needed more than polished tone: consequential answers needed sources, explicit limits, and a clear reason the response applied to them.

What Jupiter locked

Final delivery

Specified output

Answer templates

Identical questions cannot drift if the answer is a benefits decision. I worked with Data on formats, golden examples, and do/do-not rules so the model had a target, not a vibe.

Brand

Tone is not a generic assistant

ChatGPT and Gemini have a voice. Nayya needed its own. Workshops produced principles that went into the prompt system so the generated answer stayed on-brand.

What chat-only missed

Density needs a surface

Playtests showed numbers, plan comparisons, and evidence getting lost in the transcript. That was the learning. The enrollment-style surface was designed in Titan, not shipped as Jupiter.

Probed, not the product

Acting with permission

We also tested whether the system could complete a bounded task, such as an HSA contribution, with an explicit confirm. That is an agency question. It is not what Jupiter launched.

05 / The pivot

Jupiter proved chat-only was the wrong product form. Surfaces was the next system to build.

Jupiter was validated evidence for Titan’s direction, not a blind thesis. Market testing and playtest feedback showed why Surfaces and UI had to sit alongside chat: density, evidence, and deterministic work could not stay inside the transcript alone.

Jupiter

Chat as the product

Conversation handled discovery, explanation, context, and task progression inside one transcript.

→
What Titan had to build

Chat plus Surfaces

Jupiter did not ship the enrollment canvas. It produced the requirement: conversation stays primary, structured UI takes density, evidence, and work that needs a deterministic path.

Jupiter succeeded because it gave us evidence strong enough to change the thesis.