top of page

Can you blame the chatbot? The case for slow AI in healthcare

2 days ago
6 min read

Can a company blame its own chatbot? No.


In February 2024, an airline told a Canadian tribunal that its chatbot on its website was responsible for its own actions. The chatbot had given a grieving customer the wrong fare advice; the tribunal held the airline liable and ordered it to pay the difference.


It’s a small anecdotal case that made a pretty significant point clear: when AI says something on your behalf, you’re the one who said it.


That point gets heavier every quarter, because AI isn’t going anywhere and we need to accept this. The American Medical Association found that 81 percent of physicians used AI professionally in 2026, up from 38 percent in 2023. The same shift runs through healthcare marketing and sales, where AI now drafts the follow-up email, creates social media content, and writes the fourth version of the landing page.

Bar chart of physicians using AI professionally in AMA surveys: 38 percent in 2023, 66 percent in 2024 and 81 percent in 2026, more than double the 2023 rate.

Most of that AI is judged by how much it produces to keep up with the spiraling need for more and more content, which was fine until it wasn’t anymore.


In healthcare, volume eventually runs into reviewers, legal, and regulatory, for example. Most healthcare brands present to one seat and get reviewed by six, and those six people read vendor material like a regulator reads a label. Sooner or later one of them asks where a specific sentence came from, and “the model wrote it” doesn’t survive the question.


Healthcare already has its own version of the Canadian chatbot case. In 2024, the Texas attorney general settled with a clinical AI company over allegations that it overstated its accuracy to hospitals. The company denied this and agreed to explain how any accuracy figure in its marketing is defined, or have an independent auditor back it up.


The universal remedy for such instances is traceability; in this case it arrived the hard way.



The catch


The obvious fix is disclosure: tell people when AI was involved and let them judge, but the research on that is less comfortable than you’d expect.


Findings from 13 experiments by Schilke and Reimann: people who disclosed AI use were trusted less, people exposed by someone else lost more trust and the penalty shrank but never reached zero when the AI was seen as accurate.

Across 13 experiments, University of Arizona researchers Oliver Schilke and Martin Reimann found that people who disclosed their AI use were trusted less than people who didn’t, with the caveat that being exposed by someone else cost more trust than disclosing.


One detail in their findings matters more than the headline. The trust penalty shrank, though it never disappeared, among people who saw the AI as accurate. So disclosure alone won’t cut it; it has to come with something people can check.



Slow AI in healthcare


That’s why we built against the grain. I’ve been calling it slow AI, borrowing loosely from slow food.


The point of slow food, as I understand it, is knowing where every ingredient came from and who handled it on the way to your plate. The wait is a side effect.


Fast food answers a different question, which is how much you can serve for how little. Nobody asks what went into it, and the business depends on nobody asking. A lot of marketing AI works the same way, and in most categories nobody minds.


Healthcare minds, because a reviewer has to know what went into every claim, the way a slow-food kitchen knows its farms.


The metaphor breaks in one place, and it’s the one a buyer would ask about first: Will it delay workflows? No, slow AI doesn’t mean a slow answer.


For example: a committee-level brand diagnostic used to take a quarter and countless research hours, with interviews to schedule, surveys to field, and a findings deck at the end, potentially stale the moment it’s presented. Reading at that scale is what AI does well, better than we as human operators ever could, so nunNEO hands the quarter back.

Comparison of AI built for volume, which asks how much it can produce, with AI built for traceability, which asks where a claim came from.

The slow part is how each finding gets verified. The quarter you get back goes to the decisions only people can make.


What slow looks like


nunNEO.prob3, the diagnostic we build at STR3.AI, reads how a healthcare brand lands with the people who approve the purchase. Every finding gets checked against its evidence, and every report can be traced to the version of the method that produced it.


Each report says on the document itself that AI produced it. That’s the disclosure, and the rest is the evidence.


Before a report goes out, its findings are tested by a check built to refute them rather than confirm them. When a problem turns up that can’t be fixed provably, the report holds, and holding is the default. A person has to release a held report, and the record shows who released it and when.


Scores come from arithmetic. The model writes the explanation of a number and never the number itself.


The report also names what the method couldn’t see, and where the evidence was thin, it says so instead of filling the gap. That’s the rule I’d defend hardest, because a confident paragraph written over thin evidence is what volume AI produces best.


Nutrition-label-style Report Facts card for a nunNEO diagnostic report: produced by AI and stated on the report, findings tested by a check built to refute them, scores computed by arithmetic, blind spots named, method version traceable, held reports released by a person on the record and 0 percent client data used for training.

The method carries a build record for every release, and a finding that turns out to be wrong goes through a defined correction path instead of a quiet rerun. That discipline is borrowed from medical device quality practice.

I spent years in marketing and commercial roles in implant systems, digital dentistry, and dental imaging, where a product without a record of how it was made doesn’t ship. Borrowing the discipline is not a regulatory claim, and nunNEO is not a regulated device, but it does arm nunNEO with evidence-based traceability.


Client data is never used to train AI models, and our model provider doesn’t use it for training either. Every company that processes data within the ecosystem is named with its purpose in our security brief at str3.com/security. Our full statement on where AI appears in our work is at str3.com/ai.



The price


Slow has a price. Holding by default means some reports wait for a person, and we decided that was the right cost.


It does cost us the easy claims. Why? We don’t hold external certifications, so we don’t claim them, and our security brief says which controls aren’t settled yet.


It also means publishing your own mistakes. The first version of that brief contained two claims its own standard wouldn’t accept. We corrected both in the next version and said so on the page, because I’d rather publish a correction than let a transparency document quietly improve itself.


For a healthcare marketing team, the price is giving up volume as the scoreboard. Output is easy to count. Traceability is harder to count and harder to fake.



Why AI stays scaffolding


Scaffolding is the right picture for where AI belongs in this work. It lets people reach what they couldn’t reach on their own, and everyone on the site can see it and inspect it, and nobody mistakes it for the building.


Line drawing of scaffolding in front of a building, beside the text: Take the AI out of the picture. Does the finding still stand on its evidence?

The diagnostic works that way. It reads and reports but doesn’t write a brand’s claims, and people build on what it finds.


Transparency is what keeps AI in that role. Once people can’t see it, AI stops being scaffolding and becomes the building by default, and nobody signed off on that change.


The test works on any AI near your brand: does a finding still stand on its evidence once you take the AI out of the picture? If it only holds up because an AI said it, it was never a finding.


The airline’s chatbot wasn’t responsible for its own actions, and no AI near your brand will be either.


Plenty of AI will keep producing more marketing, and some of it will be really good. In healthcare, every extra piece makes one question more urgent for the people who sign off on a purchase, which is where a claim came from. I’d rather build the AI that can answer it.



What is slow AI? AI built for traceability rather than volume: every finding is checked against its evidence, and every report can be traced to the version of the method that produced it. The slow part is how each finding gets made.


Is client data used to train AI models? No. Client data is never used to train AI models, and STR3's model provider doesn't use it for training either. Every company that processes client data is named, with its purpose, at str3.com/security.


Do you use AI to write the report? Yes. nunNEO reports are generated by AI, and each report says so on the document itself. Scores come from arithmetic, and a report that fails its checks is held until a person releases it, on the record.


Drafted with AI assistance. Written, reviewed, edited, and approved by Philipp Striebe.



Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page