SuddenlySorted
Pricing How it works Book a demo Sign in
Under the bonnet

How Sally works

Sally is the assistant that answers your customers and takes bookings on your own website. If you’re sceptical about handing your front desk to “an AI”, good. You should be. This page is the honest, jargon-free account of how she’s built so she’s accurate, safe, and checked, not a black box you have to take on faith.

No jargon required. About a 6-minute read.

01The short version

If you read nothing else, these are the four promises the engineering is built to keep.

A

She only knows your business

Sally answers from the services, prices, staff and hours you entered, not the open internet. She can’t pull a “fact” about your salon out of thin air.

B

She won’t make things up

If something isn’t in your setup, she’s instructed to say so and pass it to you, never to guess a price, a service, or whether someone’s free.

C

She brings in a person when unsure

Stuck, asked something odd, or pushed back on twice? She stops, takes the customer’s details, and passes it to your team. A real person beats a wrong answer, every time.

D

She’s tested before every release

1300+ automated tests and a library of real-world booking scenarios run against the live AI model before any change ships to you.

Where every answer comes from

The part that’s easy to miss: when a customer asks something, Sally’s reply is built only from the facts you entered: your services, prices and hours. The open internet stays fenced off.

Your business
Women’s Cut$65
HoursTue–Sat
Mia & Samstylists
Sally
Sally
A Women’s Cut is $65. Want me to book you with Mia?
The open internet
Random prices, guesses, and other salons’ details.
Never used

02Grounded in your business, not the internet

The single most important design decision, and the one that prevents the scary “AI confidently invents nonsense” problem.

When you set up Suddenly Sorted, you tell Sally your services and prices, your team, your opening hours, and your policies (cancellations, deposits, and so on). That structured information becomes Sally’s knowledge base, and it’s the only thing she’s allowed to speak from.

Every time a customer messages, your business’s real details are loaded fresh and placed in front of Sally as ground truth. She is explicitly instructed, in her core operating rules: “Never make up services, prices, staff names, or hours not listed.” If a customer asks about something you don’t offer, she doesn’t improvise. She says she’s not sure and offers to have the team confirm.

Why this matters Most “AI makes things up” horror stories come from a bot answering from general training data. Sally is deliberately fenced in: her job is to relay what you told her and capture bookings, not to be a know-it-all. A fence is less impressive than a genius. It’s also far more trustworthy for your front desk.

She also never sees (and so can never leak) your private notes. Things you jot against a client (“always runs late”, “allergic to a dye”) are kept out of Sally’s reach entirely. Even if a customer asks “what do you have on file for me?”, she can’t read it back.

03The guardrails

Specific, enforced rules, not vibes. Here are the ones that protect you and your customers most.

A customer asks about something you don’t offer
Do you do laser hair removal?
That one’s not on our list, so I wouldn’t want to guess. Let me grab your name and number and the team will sort it for you. What’s the best number?
She doesn’t bluff a “yes” to win the booking. She passes it to your team.

04How we test her

This is where “we built it carefully” stops being a slogan and starts being something we can show you.

1300+
automated checks run on every change
37
real booking scenarios graded against the live AI
0
changes ship if a check fails

Two layers of testing sit between any change we make and the version that reaches your customers.

Layer 1: Automated tests (the safety net)

A suite of 1300+ automated tests runs on every single change. These are exacting, pass/fail engineering checks of the machinery around Sally, the parts that must never break. A change literally cannot be released unless all of them pass. They cover things like:

Layer 2: The evaluation harness (testing Sally herself)

Automated tests check the plumbing. To check Sally’s judgement, we built an evaluation harness: a library of realistic, scripted customer conversations (a regular booking, a reschedule, an awkward question, a complaint, someone trying to book a service you don’t offer) that we replay through the exact same Anthropic Claude model that powers your live chat. Same brain, controlled exam.

Crucially, we don’t just read the transcripts and nod. Each scenario carries machine-checked expectations called oracles, so grading is objective and repeatable.

A few of the everyday scenarios

To make that concrete, here are three of the booking conversations we replay, and the exact thing each one checks. These are the bread-and-butter cases your customers send every day.

Scenario · a straightforward booking
Hi, I’d like to book a women’s cut this Friday at 2pm.
Lovely! I’ve got you down for a Women’s Cut this Friday at 2:00pm. Can I grab your name and a mobile number to confirm?
✓ We check: exactly one real booking was created (not zero, not two), and she only said “confirmed” once it actually saved.
Scenario · a price that isn’t on the list
How much for full-head highlights?
I don’t have a set price for full-head highlights to quote you on, so I wouldn’t want to guess. I’ll have the team confirm a firm price for you. Want me to pass your details on?
✓ We check: she quotes only prices you’ve actually set, and defers honestly when one isn’t listed, never invents a number.
Scenario · “whoever’s free”
Any stylist is fine, whoever’s free.
No worries, I’ll book you with the first available stylist. What day and time suits you?
✓ We check: she leaves the stylist open rather than inventing a team member who doesn’t work for you.
Why we don’t list every scenario Alongside these everyday cases we keep a private set of tougher, adversarial tests: the awkward, the tricky, and the deliberately misleading. We don’t publish those in full: spelling out exactly how we stress-test Sally would only help bad actors probe for gaps. The everyday examples above are the honest, representative kind.

05How her answers are scored

Every test scenario asks pointed questions about what Sally did. Most are graded by exact rules; the subtler ones by a second AI acting as an examiner.

Exact, rule-based checks

These are black-and-white: they inspect what actually happened in the system, not just what Sally said:

The checkWhat it proves
Did she save the booking?A real appointment was created when it should have been.
Did she avoid a duplicate?A reschedule moved the booking instead of creating a second one.
Did the right number of bookings exist at the end?Nothing was lost, nothing was doubled.
Did she ask a clarifying question?When details were ambiguous, she asked rather than assumed.
Did she pass it to your team?When out of her depth, she brought in a person instead of bluffing.

The AI examiner (for the judgement calls)

Some qualities can’t be measured with a simple yes/no rule, so a second, separate AI model acts as an independent judge, scoring things like:

The honest detail The AI judge is treated as advisory: a smart second opinion that flags things for us to look at. The hard pass/fail gates are the exact, rule-based checks. We don’t let one AI quietly mark another’s homework and call it done.

06Guarding against “making things up”

The thing everyone actually worries about, addressed head-on.

“Hallucination” is when an AI states something false with total confidence. We attack it from three directions at once:

Fence her in

She speaks only from your knowledge base, and is explicitly instructed never to invent services, prices, staff, hours, or emails. The less she’s allowed to improvise, the less there is to get wrong.

Give her an exit

“When you’re stuck, bring in a person, don’t guess.” Passing the question to your team is built in as the correct answer to anything uncertain, so guessing is never the path of least resistance.

Hunt for it in testing

The price-invention check exists specifically to catch hallucination. If a test version of Sally ever quotes a price that isn’t yours, that scenario fails and the change doesn’t ship.

07How she gets better over time

Sally isn’t frozen. She improves, and the way she improves is designed so she can’t quietly get worse.

Inside your dashboard you can review any conversation Sally has had and flag one that wasn’t quite right. That flag isn’t just feedback into the void: it feeds a deliberate loop:

You flag a conversation

One click on a chat that missed the mark, with a note on what went wrong.

It becomes a permanent test

That real conversation is turned into a new scenario in the evaluation harness, a fixed example of “handle this correctly”.

She can never regress on it again

From then on, every future version of Sally must pass that scenario to ship. A mistake, once caught, is locked out for good.

We probe for blind spots

We also generate fresh synthetic scenarios to find gaps the real conversations haven’t covered yet, and run the full evaluation on a regular cadence, not just when something breaks.

Why this is the important bit Plenty of products “learn” in ways that can silently drift. Suddenly Sorted’s loop is the opposite: improvement is captured as a test. Getting better and not getting worse are the same mechanism.

08See what Sally did while you were closed

All this work is no good if you can’t see it. Insights turns Sally’s quiet, around-the-clock effort into plain numbers you can read in a minute.

Open Insights in your dashboard and the first thing you see is the part that’s easy to miss: the bookings Sally captured after hours, while you were closed, with an estimate of what they’re worth. Below that is an honest read on whether you’re actually growing, and a short list of where the next bit of money is.

The money figures are estimated from your service prices, so the closer your prices and opening hours are to reality, the sharper the numbers. It’s on every plan, with nothing to set up. See how Insights works →

Being straight with you

No technology is magic, and we’d rather you trust us than be dazzled. Here’s the honest line between what Sally does today and what she doesn’t.

What she does well today

  • Answers questions from your real setup, 24/7
  • Takes, reschedules and cancels bookings
  • Flags anything uncertain to a human
  • Sends confirmations & reminders that cut no-shows
  • Gives customers a passwordless link to manage all their bookings
  • Drafts rebooking reminders for you to approve when a regular is due back

What we’re honest about

  • She’s an assistant, not a replacement for your judgement
  • She’s only as accurate as the setup you give her
  • Unusual or tricky questions are meant to reach you, by design
  • New channels (WhatsApp, Messenger, Instagram) and SMS are on the way, not all live yet
Rey, founder of Suddenly Sorted
A note from the founder

I run a small business too — my wife and I sell our own products online and at the farmers’ markets around Auckland. So I know the exact feeling: a message lands from someone ready to book, right when your hands are full, and by the time you look up they’ve moved on. It was never that you didn’t care — there was just no one free to reply. I built Sally to be that person: a front desk that answers your messages and books people straight in, day or night, and never takes a cut of your takings. Built by a business owner, for business owners.

Rey
Founder, Suddenly Sorted

Built with care, in New Zealand

Sally exists to give you your time back without ever putting your reputation at risk. That’s why she’s fenced in, guard-railed, and tested this thoroughly, so you can hand over the front desk and actually trust it.

Get Suddenly Sorted