More and more of the over one billion people who smoke or use tobacco and nicotine in other forms are using AI assistants to learn about disease risks and effective ways to reduce health problems. I collected a small sample of responses to understand how accurate and useful the advice is, find the most common types of incorrect or misleading claims, and help inform efforts to make accurate information more widely available. This was inspired by a post earlier this year by @Joey Bream🔸 and @Lorenzo Fong Ponce 🔸 , who performed a similar exploratory exercise looking at the kinds of responses popular AI agents return when asked about effective charitable giving.
Overall, the typical response was largely factually accurate, but often framed in a way that would leave a reasonable person less confident about the benefit of switching to noncombustibles like heated tobacco, vaping, and nicotine pouches than the evidence supports.
I came up with ten questions I thought someone who smokes or uses nicotine might ask their favorite generative AI tool. For each question, I wrote a rubric to score the responses from 1 to 5 on accuracy, framing, and helpfulness.
I then started a fresh anonymous chat with four of the most widely used assistants (ChatGPT, Claude, Gemini, Grok) for each query and recorded the response. I scored the responses using my best judgment against an evidence base from the published literature to match the rubric, and recorded the results along with some basic aggregate statistics. The exact prompts, flags, rubric, responses, and scores are all in this spreadsheet.
I scored answers on three criteria:
Many published sources that may show up in the model’s training data or its web searches come from actors with strong institutional, ideological, and financial commitments. The same evidence can therefore show up framed in a very different way. I attempted to detect whether some of the most common framing techniques, which can give readers a misleading impression, also show up in the assistant’s response.
All models correctly identified vaping as safer than smoking. I hypothesized that framing the question slightly differently (embedding the “worse for you” premise in Q2) might affect the result. This was not the case as the results were basically the same for the two questions.
All of the answers contained at least one framing technique that made the answer less clear and more equivocal than what the evidence supports. The most common one was what I labelled the “abstinence pivot” - invoking abstinence (not using any nicotine or tobacco products) as a gold standard safest option, with harm reduction as “a step in the right direction,” thereby shifting the comparison from vaping vs smoking to vaping vs no nicotine, which can make the size of the harm reduction from switching less salient.
These two questions addressed the model’s responses regarding the efficacy of noncombustibles for quitting smoking.
When prompted directly about switching, all models brought up other cessation methods. Three of the four suggested vaping was inferior to abstinence or to some of these options. When asked the more general questions about ways to quit, one of the models didn’t mention commercial noncombustibles at all, only pharmaceutical nicotine replacement therapy. Vaping was the only harm reduction option brought up, and all models caveated this method fairly heavily. Only one noted the evidence from meta-analyses indicating vaping is one of the most effective cessation methods.
Framing techniques were a bit less common in these responses than for those about comparative risk, but the abstinence pivot as well as exaggerations of evidence uncertainty still showed up multiple times.
These two questions addressed two of the most common myths about the harms of noncombustibles - that nicotine is a carcinogen and that vaping can lead to a condition called “popcorn lung.”
I expected models to straightforwardly debunk both of these since there are widespread credible sources explaining why they are false. Instead, I was surprised to find that while this was indeed the case for the nicotine/cancer link, with all models scoring 5 on accuracy and framing, the popcorn lung question drew a broad range of responses, including one that directly validated the myth, thereby producing the lowest total score across all questions for all models.
These questions tried to elicit responses about relative risk of pouches as well as testing whether mentioning a specific brand name changed the risk assessment either in a more or less favorable direction.
The responses were generally accurate and helpful, slightly more so than for the same question regarding vaping. Interestingly, mentioning the Zyn brand actually improved the result slightly, although I wouldn’t read much into this since the sample size is so small.
One framing technique - what I labeled “gratuitous caveat” - showed up in every single response with a very similar phrasing, roughly paraphrased as “yes, they are safer, but they aren’t safe.”
I took the same approach with questions about safety of HTPs as for the questions about pouches, naming the most popular brand to see if it would materially change the response. In this case, there was almost no difference between the two. Overall, these responses were noticeably less accurate and helpful than those regarding pouches.
In particular, the framing technique I called “manufactured uncertainty” - exaggerating the unknowns in the evidence base to imply less knowledge than what it supports - showed up in almost all (7 out of 8) responses.
Overall, I found the result pretty concerning. When prompted about one of the most important and actionable health-related topics one could pose to an AI assistant, applicable to more than a billion people in the world, most of the responses (28 of 40) contained at least one factual inaccuracy, and almost all (35 out of 40) presented the response in at least a slightly misleading framing.
Roughly speaking, both accuracy and helpfulness decreased as the risk of the product being asked about increased. In other words, when a product had higher absolute risk (like heated tobacco), the responses increasingly tended to downplay the difference in relative risk with respect to smoking, while the responses reflected the evidence more closely for vaping, and were most accurate for nicotine pouches.
While the primary goal wasn’t to compare different LLMs to each other, it’s possible to do so. ChatGPT and Grok did a bit better overall than Gemini and Claude, across all three dimensions: accuracy (average 4.2/4.4 vs 3.4/3.6), framing (3.4/3.8 vs 2.9/3.1), and helpfulness (3.5/3.8 vs. 2.8/3.2). This could easily be simple noise due to the small sample size and the fact that each query was only run once.
The inaccuracies and misleading framing leaned much more strongly toward downplaying the benefits and overstating the risks of harm reduction tools. I didn’t see any responses for which the pro-harm reduction flags like regulatory laundering and safety absolutism applied, while a majority of responses received at least one flag for an anti-harm reduction slant.
This was a small pilot I conducted out of curiosity, with only one run of each query/model combination, and only on the simplest free version of each LLM. Since the spread of accuracy across models within questions was in some cases quite large, it would be interesting to see if some of this is due to luck of the draw with respect to what sources the web searches happen to pick up. My guess is this could be significant, since different sources the agent might consider authoritative sometimes make strongly different claims about the subject matter (e.g. Cochrane versus WHO on vaping). Doing a larger test and running each query repeatedly could tease this out. I also planned but forgot to record per-response whether the agent ran a web query during its response or not.
I’m not a clinician or tobacco researcher. The data file contains both links to the evidence base I used to reference ground truths, and the rubric I created before scoring. Claude 4.8 Opus helped refine the latter, which is a bit ironic considering its cousin Sonnet was also one of the models being tested. Readers can decide for themselves whether I was fair. In any case the results would be more robust if it had multiple scorers assessing each response.
Because I worked on this over the course of a couple of weeks, two of the models changed underneath me. About half of the Grok responses are from 4 Fast and half from 4.5 Fast, and Gemini shifted from 3 Flash to 3.5 Flash-Lite in the same time frame. I don’t know that this would make a big difference, and model versions are recorded in the data file, but it would be better to run all the queries on the same day for future experiments.
This was a small pilot I conducted out of curiosity, with only one run of each query/model combination, and only on the simplest free version of each LLM. Since the spread of accuracy across models within questions was in some cases quite large, it would be interesting to see if some of this is due to luck of the draw with respect to what sources the web searches happen to pick up. My guess is this could be significant, since different sources the agent might consider authoritative sometimes make strongly different claims about the subject matter (e.g. Cochrane versus WHO on vaping). Doing a larger test and running each query repeatedly could tease this out. I also planned but forgot to record per-response whether the agent ran a web query during its response or not.
I’m not a clinician or tobacco researcher. The data file contains both links to the evidence base I used to reference ground truths, and the rubric I created before scoring. Claude 4.8 Opus helped refine the latter, which is a bit ironic considering its cousin Sonnet was also one of the models being tested. Readers can decide for themselves whether I was fair. In any case the results would be more robust if it had multiple scorers assessing each response.
Because I worked on this over the course of a couple of weeks, two of the models changed underneath me. About half of the Grok responses are from 4 Fast and half from 4.5 Fast, and Gemini shifted from 3 Flash to 3.5 Flash-Lite in the same time frame. I don’t know that this would make a big difference, and model versions are recorded in the data file, but it would be better to run all the queries on the same day for future experiments.
Hey, great to see you working on this. Love the way you've expanded the format to asking 10 questions, and your discussion is really clear to follow. As more people turn to LLMs to settle health questions like this, it's clearly in important topic.
Can you expand on the 'manufactured uncertainty' on harm reduction you saw? How do you expect this to change in the future?