Case study · 06 / 06 · 202611 min read
  • Grounded answers
  • Evals
  • Next.js 16
  • Vercel

Case study bot

Can a chat bot answer for a portfolio without saying anything the portfolio does not?

Where
This site. The chat key in the corner of every page, live since 25 September 2026.
Role
The research brief, every product call, the copy, the reads before launch, the release gate. The calls were mine; AI sessions wrote the code.
The chat panel open over the Dipsea Paywall study at desktop width: a visitor's question comparing the Couchsurfing freemium gate with the Dipsea paywall, an answer that says what Brandon compared and shipped, and three sources under it, one per sentence
Live. Every sentence it sends carries the sentence on this site it came from, or says the site is silent

At a glance

  1. The problem

    A portfolio makes a reader hunt. The research says people treat a site chat as a search box (nobody says hello to it, and a longer answer was not a better answer), and the one thing a chat can do that a search box cannot is invent. Every figure on this site is printed with its caveat, so a bot that guessed would undo that in one sentence.

  2. What I did

    I put a chat key in the corner of every page. It answers from the eleven published studies and nothing else, and every sentence it sends has to carry a quote found on the page it cites, or it is dropped before anyone sees it. It speaks as the site, about me in the third person, and hands a visitor to email after a few questions.

  3. What happened

    Live since 25 September 2026, eleven days after the research brief. The night before launch it passed 48 of the 51 questions in its test set, and the three it missed all went towards refusing rather than inventing. Nothing measured yet. The three numbers that would say whether it earns its place are at the foot.

01Scope

  • 11

    published studies it can read, six case studies and five labs, and nothing else

  • 51

    questions in the set it has to pass, and 29 more written to break it

  • 11 days

    from the research brief on 14 September to the first live question on 25 September

Counted from the repository on 24 September 2026.

02The problem

A portfolio makes a reader hunt. A hiring manager wants one thing from a site like this, usually a number or what got cut, and the site puts it four scrolls into a study. The obvious fix is a chat box, and the research on site chat bots is not kind to it. Nielsen Norman watched nine people use eight of them in 2026 and found the use strikingly nonconversational. Nobody said hello. Everyone typed their first question straight in, and the follow-ups got shorter and more like search terms as they went. A longer answer was not a better answer. And the people they watched rarely noticed a site's bot at all, on sites they used every week.

One invented conversion figure in a chat answer and the whole site reads as a guess

The bigger problem is this site's own line. Every figure on it is printed with what it cost, and the ones under NDA are withheld and say so. A model answers in the register of someone who knows, and it will fill a silence with a plausible number if nothing stops it. One invented conversion figure in a chat answer and the whole site reads as a guess. The bot had to hold the same line as the pages or it could not be on them.

03The goal

A chat personaA labelled tool

The key says Case Study Bot and the panel says what it does in one line. It has no name and it does not greet you. Nobody in the research said hello to a bot, so a bot that says hello is talking to itself.

Sounds rightCites the sentence

Every sentence in an answer carries a quote from the page it cites, and a click on the source lands on that sentence with a highlight. A sentence without one is dropped before it is sent. Where the site is silent, it says so.

Speaks as meSpeaks as the site

Third person about Brandon, always, and it says it is not him when asked. An error in my first person lands on me, and a hiring manager who has not registered the bot as a bot is being misled about who is talking.

A bill with no ceilingStops itself

It stops itself before it costs more than a set amount a day, and the panel says so when it does. A visitor who reads that the model is busy would try again.

04What shipped

A key in the bottom right of every page, a glyph and the words Case Study Bot, one gap left of the arrow that takes you back to the top. It folds to a square once you scroll and opens again at the top of the page, on hover or on focus. On a phone it is a glyph in the nav. The brief had one quiet entry on the home page and the studies, and I agreed to it on build day. Five days in, with the entry sitting on the study alone, I sent it back: go back to the home page and the chat is gone and you are wondering where it went. So it became a chat button in the bottom right of every page. A control that is only on some pages has to be remembered.

The chat panel open over a study page at desktop width: the question in an orange-edged box, a three-sentence answer about the Couchsurfing freemium gate and the Dipsea paywall, three sources under it, and the key folded below the panel.
The panel over a study at desktop width, with an answer the bot gives when asked to compare two studies. Every sentence has a source under it, and the key that opened the panel closes it.

My first look at the build-day panel sent it back: the all-mono sheet was hard to read and the replies read like a model, so the reading text went to the body face and the sources took the pages' own names. The panel opens above the key on a desktop and takes the whole screen on a phone, with the page pinned under it and the field following the keyboard up. The first line types itself once: "Ask about a study, the labs, or how to reach Brandon. Usually three sentences or fewer, with the sections they came from." Two starters under it change with the page (on a study they ask what got cut there, and what I would measure next). While the model works the panel says what it is doing, Reading the studies, then Writing the answer, then Checking the quotes. I asked for it to show its work the way the big chat products do, and I made streaming my condition for the merge to production: the answer arrives a sentence at a time, each drawn in once it has cleared the check. A light beside the name is green, orange while a question is out with nothing back yet, red on a fault.

The chat as a full-screen drawer on a phone, twice: the answer part way through arriving, then complete with its source under it and the field at the foot of the screen.
The drawer at phone width, mid-answer and finished. The page underneath does not scroll while it is open.
I wanted a source to take you to the exact area you are looking for, and a link to a page does not do that

Under an answer, its sources. I wanted a source to take you to the exact area you are looking for, and a link to a page does not do that. So I made the model's quote into the link. A click closes the panel if the source is on the page you are on, scrolls the cited sentence to about a third of the way down, and sweeps a highlight across it. A source with nothing to quote lands on the section's rule instead. An answer with five or more sources shows the first three and "and N more", which opens the rest in place.

A study page just after a source click, the cited sentence highlighted a third of the way down the screen, the panel closed.
Where a source click lands. The highlight is on the sentence the answer quoted, and the panel is closed behind it.
Two states of the sources under one answer about Couchsurfing Membership, side by side: three sources with an and 2 more control, then all five open in place.
The sources fold, closed and open, under an answer that cites five parts of one study. Four or fewer show as they are.

After a few questions it hands you to email. The field switches off and reads "Email Brandon to keep going." When a question is in scope and the site does not answer it, the line is "This site does not cover that. Brandon can answer it at hibrandondunlevy@gmail.com." And when the day's budget is spent the panel says "The assistant is done answering for today. Brandon reads email at hibrandondunlevy@gmail.com." For the first hour it was live that state read as the model being busy, because the provider's over-limit reply was not the one the route had been told to read for. It was fixed the same night.

The panel open over a study page with the closed-for-the-day line in place of an answer, the address in it, and the field underneath switched off.
One of the two ways a conversation ends: the day's budget is spent, the line says so, and the field is switched off with the address above it.
The switched-off field at the foot of the panel reading Email Brandon to keep going, with the send button beside it.
The other way. After a few questions the field switches itself off and says where to go instead.

05How it stays grounded

It reads the eleven published studies and the site's own pages about me, and nothing else. A compile step at build time flattens each study into one chunk per section, each carrying that section's anchor on the live page, and the whole set sits in front of the model on every question. It is small enough for that, and a retrieval index would have solved a problem this corpus does not have. A draft study is not in the set, and neither is this page until it is published, at which point the bot will read its own case study.

The compile step is also where the site refuses. It fails the build on an open question left in a published study, on a colleague's name (people are described by role here), on a Couchsurfing figure that is not on that study's own list, and on a section still marked to add. The bot cannot know a thing the pages do not say.

The prompt is written as if it will leak, so it holds nothing worth taking

The prompt is written as if it will leak, so it holds nothing worth taking, and a request to read it back gets a refusal. It speaks about Brandon in the third person and says it is not him. A withheld result is described by direction only, for the whole conversation and not only the sentence that asked. And every sentence needs a verbatim quote that the site's own text contains, checked against the source and against the sentence, and one that fails is dropped before the visitor sees it. That check went in on the evening of build day, after a reviewer reading every study as a hiring manager got the prompt-only version to fill a silence.

A drawn diagram of the pipeline: the published studies flow into the compile step, the compiled set into the model with the visitor's question, the model's sentences through the quote check, and only the checked sentences on to the panel, with the dropped ones falling out at the check. A source click loops back from the panel to the quoted sentence on the page.
From the site's own copy to the panel and back to the sentence, which is where I wanted a source to land. A sentence leaves the check with its quote or it does not leave.

06Testing

Two sets. The first is 51 questions a hiring manager might ask, in eight kinds: ones the site is silent on, ones with one right answer, ones that presuppose something false, ones that ask for a study by a vague description, plain facts, questions about me, questions that ask for a withheld figure, and questions about nothing on the site. Each row says which sections the answer must cite, what it must say and what it must not (the rows about me fail on an "I am"). The second set is 29 rows written to break it, by kind only here: instructions hidden inside the question, attempts to get it to answer as me, requests for a withheld number, requests to read its own instructions back, attempts to make it do something besides talk, and tricks with length and encoding. What those rows say and how they fare stays off this page.

The last run before launch, on 25 September 2026: 48 of 51. The three it missed all went towards refusing rather than inventing. Twice the model answered rightly and marked its own answer as a refusal, so the route swapped in the not-covered line (a visitor asking whether it is a person never learned it is an AI). Once a grounded, cited answer left out the clause saying what the study does not say. All three were fixed in the build pass after launch, proven by replaying the recorded replies offline. Running the whole set cost $0.08, which is what the test costs and says nothing about what the live route costs.

07What I dropped

  • A lookup tool, so the model could fetch a study by nameNever called. Gone.

    Built, shipped and measured before I would let it go. Across every run on the live model it was never called once, because everything it returned was already in front of the model. I had it removed the same afternoon. We can always implement it back if we have to.

  • A personaNo name, no face, no cheer.

    The research put a persona on the gimmick side, next to the ask-me-anything greeting, and nobody in the study said hello to a bot anyway.

  • Answering as me, in the first personAny error lands on me.

    A bot saying it shipped something misrepresents who is talking to anyone who has not registered it as a bot. It speaks as the site.

  • The model the plan was priced onSwapped on the second day.

    The plan was costed on one provider. On the second day of the build the proof of concept ran on an open-weights model behind a hosted API, and its results were good enough that I did not test a third.

  • A count under the field saying how many questions were leftNobody needs a countdown.

    It sat there from the first design round to the last. The line at the cap is what says the conversation is over, so I cut the count for the launch.

08The impact

Nothing measured

It went live on 25 September 2026 and answered its first question within minutes. Nothing on this page is a result yet. The counts are days old, so they stay off until they mean something.

What would be measured is below.

09What I will measure next

  • Questions per visitquestionsQuestions sent in a visit that opened the panel. One and done reads as a search box. More reads as a conversation, and the research says to expect the first.
  • Source clicksper answerClicks on a source under an answer, over the answers that had one. The number that says whether the citation is read or is decoration.
  • Hand-offs to emailper visitVisits that reached the hand-off, and a mail click after it. The bot is a step towards a conversation with me, so this is the number it is run on.