Published on 21 September 2026

Analysis · Addictions · Methodology

GLP-1 receptor agonists and alcohol: what holds up once trials and cohorts are kept apart

◆ Collection
Addiction Science & Clinical Practice · 2026 ; 21 : 8 · Sinha and Ghosal
DOI 10.1186/s13722-025-00637-z
PMID 41350683
Scientific 73
Editorial 78

The essentials

Nine studies, two different worlds. Three randomized trials totaling 430 patients, and six observational studies whose authors report data on 2,740,207 people. The authors refuse to pool them and analyze the two bodies of evidence separately, which is the methodological decision that defines this review. The result: in the randomized trials, the pooled estimate reaches significance for none of consumption, drinks per drinking day, or craving. In the observational studies, an association with fewer alcohol-related events, with a hazard ratio of 0.64. The authors themselves describe the observational associations as hypothesis-generating and state that the current evidence does not yet support a recommendation. That caution is warranted: more than 99.9% of the combined sample comes from non-randomized data, where the observed effect may reflect who receives the treatment as much as the treatment itself.

Context

The question usually arrives through the patient, not through the literature. People treated for diabetes or obesity report a spontaneous drop in their urge to drink, the press picks it up, and the patient in front of you asks whether the same treatment could help with alcohol. A useful answer needs to be more than an impression.

The disorder this question concerns has a problem that this news tends to obscure: the medications already approved for alcohol use disorder remain markedly underprescribed, a point the authors themselves make when they place naltrexone and acamprosate as first-line options for as long as randomized evidence for GLP-1 agonists is lacking. Discussing a molecule with no approval in this indication before offering those that already have one reverses the order of priorities.

Mechanism

A plausible hypothesis, and what the study actually tests

The hypothesis in this field holds that glucagon-like peptide-1 receptor agonists act on reward circuits shared by food intake and substance use, which would explain an effect on craving. This hypothesis is consistent with preclinical work, and it is plausible. It is not what this review measures.

Element Status in this review
Effect on self-reported consumption Measured in the randomized arm only
FindingEndpoint of the randomized arm, expressed as a standardized mean difference. The observational arm does not measure consumption itself but recorded events: incidence or recurrence of the disorder, hospitalization, or acute intoxication.
Effect on craving Measured, but imprecise
FindingOnly two trials, with a confidence interval wide enough to exclude neither a substantial benefit nor a reverse effect. This estimate supports no conclusion, in either direction.
Neurobiological pathway involved Not measured
FindingNo mechanistic intermediate endpoint, no imaging, no biomarker. The shared-reward hypothesis remains a hypothesis.
Comparison between molecules Not valid
FindingIn the randomized arm, each molecular subgroup contains a single trial. Comparing molecules against one another amounts to comparing three isolated trials, a limitation the authors state explicitly.

The study at a glance

Population
Adults aged 18 and older, with or without alcohol use disorder, often with obesity or type 2 diabetes. In the randomized trials, mean age ranges from 42 to 48, with 60 to 70% men. In the cohorts, mean age ranges from 46 to 66.
Intervention
Any GLP-1 receptor agonist: semaglutide, liraglutide, exenatide, dulaglutide, plus one cohort covering dual GIP/GLP-1 agonists, including tirzepatide.
Comparator
Placebo in all three randomized trials. In the cohorts: DPP-4 inhibitors, no exposure, obesity treatments without a GLP-1 agonist, no treatment, or usual care. Eligibility criteria also allowed naltrexone and acamprosate as comparators, but no included study used them.
Outcomes
Alcohol consumption and craving, expressed as standardized mean differences in the randomized arm. Alcohol-related events, defined as incidence or recurrence of alcohol use disorder, hospitalization for a substance use disorder, or acute alcohol intoxication, expressed as hazard ratios in the observational arm.
Design
Systematic review and meta-analysis with separate arms. Three randomized trials totaling 430 patients, and six observational studies whose authors report 2,740,207 people. Searches of PubMed, Embase, and Cochrane from January 1, 2005 to May 3, 2025. Registered with PROSPERO under CRD420251045294, reported per PRISMA 2020. Random-effects model with restricted maximum likelihood estimation and the Hartung-Knapp adjustment, sensitivity analysis using DerSimonian and Laird, trim-and-fill for the trials, Egger’s test for the cohorts, prediction intervals, risk of bias assessed with RoB 2 and the Newcastle-Ottawa scale, certainty rated with GRADE.

Results

430
Patients randomized, against a combined sample of more than 2.7 million. All of the apparent weight of this literature comes from the more than 99.9% who were not randomized, and that is also where the only significant result comes from.

The randomized arm first, since it is the one that answers the question of a treatment effect.

Outcome, randomized arm Result
Total alcohol consumption, 3 trials Standardized mean difference 0.24 favoring treatment, 95% confidence interval 0.70 favoring treatment to 0.23 against, p = 0.159, I² = 54.1%
ReadingThe interval includes the null value. It also includes a moderate benefit. This result does not demonstrate an absence of effect, it shows that three trials and 430 patients are not enough to establish one. Taken individually, two of the three trials nonetheless report a significant effect: it is the pooling, with the Hartung-Knapp adjustment, that widens the interval.
Drinks per drinking day, 3 trials Standardized mean difference 0.23 favoring treatment, 95% confidence interval 0.64 favoring treatment to 0.19 against, p = 0.169
ReadingSame direction, same lack of significance. The authors note that the standard drink is not defined the same way across countries, 14 g in the United States and 12 g in Denmark, and that no conversion into grams per day was possible. Comparability of volumes across trials is therefore imperfect.
Craving, 2 trials Standardized mean difference 0.14 favoring treatment, 95% confidence interval 2.84 favoring treatment to 2.55 against, I² = 82.9%
ReadingThis interval is uninterpretable. It covers nearly the entire range a standardized mean difference can take. Combined with heterogeneity of 82.9%, it supports no claim, positive or negative.
Heavy-drinking days Odds ratio 0.84, 95% confidence interval 0.71 to 1.00, one trial, not pooled
ReadingThis outcome was reported by only one trial, the one testing semaglutide. It was therefore not meta-analyzed and was not rated by GRADE. There is no pooled result for this outcome, which is not the same as a null pooled result.
GRADE certainty, randomized arm Low for consumption and for drinks per day, very low for craving
ReadingRating assigned by the authors in their supplementary material, consistent with the very small number of trials and the imprecision of the estimates.

The observational arm next, which carries the only significant signal in the article.

Outcome, observational arm Result
Alcohol-related events, 11 estimates Hazard ratio 0.64, 95% confidence interval 0.59 to 0.69, p < 0.001, I² = 19.1%
ReadingA strong, consistent association. It describes what is observed among people who receive these treatments, not what their administration would produce in a given person. Note that this headline estimate pools 11 estimates from only 6 cohorts, mixing chronic disorder, hospitalization, and acute intoxication.
Alcohol use disorder, 8 estimates Hazard ratio 0.66, 95% confidence interval 0.63 to 0.70, p < 0.0001
ReadingSame order of magnitude. The authors report no heterogeneity, with a prediction interval of 0.61 to 0.72. Consistency across cohorts does not correct a bias common to all of them.
Substance use disorder, 2 estimates Hazard ratio 0.66, 95% confidence interval 0.18 to 2.48, p = 0.156
ReadingA similar point estimate, incomparable precision. The same number can be solid or empty depending on its interval, and that is exactly what this row shows. Both estimates come from a single Swedish cohort.
Alcohol intoxications Hazard ratio 0.50, 95% confidence interval 0.42 to 0.60, single cohort
ReadingA result from a single study, on dual GIP/GLP-1 agonists. It illustrates the signal, it does not confirm it. The authors treat it narratively while still including it in the headline estimate of 0.64.
GRADE certainty, observational arm Moderate
ReadingRating assigned by the authors. Moderate certainty on observational data does not convert into evidence of efficacy.

A signal favoring semaglutide on craving is reported with a subgroup-difference test at p = 0.024, in the randomized arm. Each molecular subgroup contains a single study. This value should not be treated as an established difference between molecules. The same reasoning applies to the observational arm, where the difference between molecules is reported at p < 0.001 even though three of the four subgroups rest on a single study, a limitation the authors themselves acknowledge in their discussion.

Quality control

Point checked Verdict
Separation of designs Strict between arms
FindingRefusal to pool randomized trials and observational data, distinct effect measures, separate analyses for alcohol and for substance use. This is the article’s strongest point. Separation is less strict within the observational arm itself, where the headline figure of 0.64 still combines chronic disorder, hospitalization, and acute intoxication.
Registered protocol Yes
FindingPROSPERO registration under CRD420251045294, reported per PRISMA 2020.
Statistical method Suited to a small number of studies
FindingHartung-Knapp adjustment, appropriate when the number of studies is small, with a consistent sensitivity analysis using DerSimonian and Laird. Prediction intervals provided.
Independence of pooled estimates Not addressed
FindingEleven estimates from six cohorts are pooled as though independent, with one study contributing four and two others contributing two each. Nothing in the text indicates a correction for within-study correlation, which tends to artificially narrow the confidence interval.
Risk-of-bias assessment Two tools
FindingRoB 2 for the trials, with an overall risk rated low to moderate, and the Newcastle-Ottawa scale for the cohorts, with scores of 7 to 8 out of 9. Separate GRADE ratings by arm.
Size of the randomized corpus Very small
FindingThree trials, 430 patients, durations of 9 to 26 weeks. Underpowered to detect a moderate effect on consumption.
Residual confounding, observational arm Uncontrolled
FindingAcknowledged by the authors, who cite socioeconomic status, psychiatric comorbidities, and lifestyle factors as unmeasured. People who receive these treatments and continue them differ from those who do not, on dimensions the models do not all capture.
Selection of observational studies Restrictive
FindingThe requirement for robust confounding control, propensity-score matching, within-individual design, or sufficiently adjusted multivariable regression, led to the exclusion of 24 studies at the full-text stage. The authors themselves note that this filter favors more recent publications and may introduce selection bias.
Internal consistency of the figures Discrepancies noted
FindingRecounting the sample sizes in Table 1 gives 2,690,207 people for the observational arm, while the text and abstract report 2,740,207, a discrepancy of 50,000. Several other values differ between the text, the abstract, and the supplementary material: heterogeneity for the main observational analysis is given as 19.1% in the text and 20.5% in the supplement, the alcohol subgroup as 0.64 (0.58 to 0.70) in the text and 0.67 (0.62 to 0.72) in the supplement, and the lower bound of the substance-use result as 0.17 in the text and 0.18 in the abstract. These discrepancies overturn no conclusion, but they call for careful reading.
Conflicts of interest and funding None declared
FindingThe authors declare no conflicts of interest. No funding source is mentioned in the published text. The authors report using an artificial intelligence tool solely for language editing, with the scientific content and statistical analysis remaining their own.
Data availability On request
FindingNo public repository; data are stated to be available on reasonable request. The software and analysis package are named, R version 4.3.1 and the meta package, but the analysis code itself is not shared.

Critical appraisal

Domain Judgment
Methodological rigor Solid
FindingThe review does what should be done when two incompatible bodies of evidence answer the same question: it keeps them apart and lets the reader see that they do not say the same thing. The protocol is registered, the bias tools match each design, and sensitivity analyses are present. The main criticism concerns execution rather than design: several figures do not match across sections, and the non-independence of the observational estimates is not addressed.
Fit between claim and evidence Measured in the body
FindingNon-significant results are presented as such, the observational associations are labeled hypothesis-generating, and the authors write that the current evidence does not yet support a recommendation. One reservation: the abstract states that semaglutide and the dual agonists have more potent effects, with a p-value, while the discussion acknowledges that this comparison is not statistically valid. This is the one sentence in the article that goes further than its data.
Randomized evidence base Insufficient
FindingThree short, small trials, only one of which measures heavy-drinking days and one of which is open-label. The pooled craving estimate, from two trials, is unusable. The field lacks trials, not commentary.
Interpretation of the observational arm A classic trap
FindingA sample of several million produces narrow confidence intervals that create an impression of certainty. That precision concerns the observed association, not the treatment effect. A confounding bias is not diluted as the sample grows, it is simply measured with more precision.

Level of evidence

Scientific73/100
Editorial78/100

Confidence is high in the overall conduct of the review and in how the authors themselves read it in their discussion. It is also high on a useful negative finding: to date, the pooled estimate from the available randomized trials does not establish a significant effect of these molecules on alcohol consumption or on craving. It is low on the causal interpretation of the observational signal, however many people it draws on.

What is demonstrated: that pooling the available randomized trials shows no effect, and that these trials are too few and too short to rule one out. It should be added that two of the three trials, taken individually, report a significant effect on consumption, which rules out presenting the randomized arm as uniformly negative. What is suggested: that an effect exists, a hypothesis made plausible by the convergence of the cohorts and by preclinical data. What belongs to expert opinion: the idea that the observational signal foreshadows what the large ongoing trials will show.

Symmetry is required here, in both directions. The lack of significance in the pooled trial estimate does not prove that these molecules have no effect on alcohol. And the size of the observational signal does not prove that they do. Both claims are wrong for the same reason.

The colleague test

What an experienced colleague might say about this study in two minutes, between two consultations.

“Two million seven hundred thousand patients observed do not replace four hundred and thirty randomized patients. The pooled trials show nothing, the cohorts show a great deal, and that is almost always the order in which these things end up deflating. I am not changing my prescribing, and I start by offering what already has an approval.”

Translated for practice: this article changes no prescription. What it does provide is a precise answer for the patient who asks the question, and a reminder of priority: the medications approved for alcohol use disorder remain underused, and that is where the accessible benefit lies today.

What you can do with this

  • Answer the patient who asks whether their diabetes or obesity treatment could help with alcohol: the signal exists, it is plausible, it is not established by the randomized trials, and it is not an approved indication for the product.
  • Do not start a GLP-1 receptor agonist with the intent of treating alcohol consumption. No regulatory approval covers this indication, and off-label prescribing carries its own obligations, which vary by jurisdiction: in France, for instance, it is subject to strict conditions, and clinicians elsewhere should check the equivalent rules where they practice.
  • In a patient already receiving this treatment for an approved indication who also has alcohol use disorder, there is no reason to stop it for that reason, nor to expect a benefit on alcohol from it.
  • Revisit the question of approved treatments for alcohol use disorder, whose prescribing rate remains well below what the literature justifies, a point the authors themselves make.
  • Use this article as a reading exercise: two bodies of evidence, two effect measures, two opposite conclusions, and a single figure circulating in the press.

Frequently asked questions

Why not pool the trials and the cohorts to gain power?

Because they do not measure the same thing. The trials estimate a standardized mean difference on consumption, the cohorts a hazard ratio on recorded events. Adding different quantities together produces a number that answers no real question. The authors themselves justify this refusal by the difference in effect measures and designs. And because several million observations would mechanically overwhelm 430 randomizations, turning the least reliable result into the final answer.

Do the negative trials prove that these molecules do not work on alcohol?

No. Three short trials totaling 430 patients lack the power to rule out a moderate effect. The confidence interval for the pooled estimate on consumption remains compatible with a real benefit, and two of the three trials individually report a significant effect. An absence of proof of effect is not proof of absence of effect.

The hazard ratio of 0.64 looks very precise, why not trust it?

Because the precision of an estimate and its validity are two different things. A narrow confidence interval means the association is measured accurately. It says nothing about whether that association reflects a treatment effect or a difference between the people who receive the treatment and those who do not. A technical factor compounds this here: the interval rests on eleven estimates drawn from only six cohorts, treated as independent.

Is semaglutide more effective than the others?

Nothing in this article supports that claim. The signal referenced comes from subgroup analyses where each molecule is represented by a single study in the randomized arm, and by a single study for three of the four categories in the observational arm. Comparing these subgroups amounts to comparing isolated studies, a limitation the authors state explicitly.

What would settle the question?

Randomized trials of sufficient size, with a validated consumption endpoint, a duration beyond a few months, and a population that actually includes patients with alcohol use disorder. The article’s supplementary material lists ten ongoing trials in this field, one already completed and one testing a molecule from a different class. This review exists precisely to hold the place until then, without deciding in place of the trials.

Annotated bibliography

Source study. Sinha B, Ghosal S. The effects of glucagon-like peptide-1 receptor agonists (GLP1-RAs) on alcohol-related outcomes: a systematic review and meta-analysis. Addiction Science & Clinical Practice, 2025 Dec 5;21(1):8. DOI 10.1186/s13722-025-00637-z. PMID 41350683. The article is dated 2026 by volume even though it went online on December 5, 2025: bibliographic registries index the year a paper goes online, the publisher cites the volume year. Two authors only, both affiliated with institutions in Kolkata, India. No funding mentioned, no conflicts of interest declared. PROSPERO protocol CRD420251045294.

The three randomized trials included. Hendershot CS, Bremner MP, Paladino MB, et al. Once-weekly semaglutide in adults with alcohol use disorder: a randomized clinical trial. JAMA Psychiatry, 2025, volume 82, issue 4, pages 395 to 405. Semaglutide versus placebo, 48 patients, 9 weeks, open-label trial. Klausen MK, Jensen ME, Møller M, et al. Exenatide once weekly for alcohol use disorder investigated in a randomized, placebo-controlled clinical trial. JCI Insight, 2022, volume 7, issue 19, article e159863. Exenatide versus placebo, 127 patients, 26 weeks. Probst L, Monnerat S, Vogt DR, et al. Effects of dulaglutide on alcohol consumption during smoking cessation. JCI Insight, 2023, volume 8, issue 22, article e170419. Dulaglutide versus placebo, 255 patients undergoing smoking cessation, 12 weeks. This last trial therefore enrolled a population recruited to quit smoking, which limits its direct transfer to an alcohol-focused addiction consultation.

The six cohorts. Wium-Andersen IK et al., Basic and Clinical Pharmacology and Toxicology, 2022, volume 131, issue 5, pages 372 to 379, Denmark. Lähteenvuo M et al., JAMA Psychiatry, volume 82, issue 1, pages 94 to 98, Sweden, within-individual design. Wang W et al., Nature Communications, 2024, volume 15, article 4548, United States. Qeadan F et al., Addiction, 2025, volume 120, issue 2, pages 236 to 250, United States. Farokhnia M et al., Journal of Clinical Investigation, 2025, volume 135, issue 9, article e188314, United Kingdom and United States. Xie Y, Choi T, Al-Aly Z, Nature Medicine, 2025, volume 31, issue 3, pages 951 to 962, United States. Worth noting: the Lähteenvuo study is cited as dating from 2024 throughout the body of the article, while its own reference entry carries the year 2025, an inconsistency in the source publication.

Competing synthesis. de Faria Moraes B et al. Impact of glucagon-like peptide-1 receptor agonists on alcohol consumption and liver-related outcomes: a systematic review and meta-analysis. Drug and Alcohol Dependence, 2025, volume 275, article 112840. This meta-analysis pools eight studies and reports a hazard ratio of 0.56, with heterogeneity of 63.6%, a more favorable reading. The authors of the review analyzed here attribute the difference to Moraes’ inclusion of liver-related endpoints: the divergence therefore also concerns the scope of the outcomes, not only the method. The two articles read well together, and comparing them is more instructive than reading either alone.

Regulatory context, to be checked against local rules. None of the molecules discussed here holds a regulatory approval for alcohol use disorder. Off-label prescribing exists in most health systems, but it is governed by its own conditions, which differ from one jurisdiction to the next. In France, for instance, prescribing outside an approved indication is subject to strict conditions, notably the absence of an appropriate approved alternative, and it is useful to keep three distinct questions apart there: whether a given product is approved for an indication, whether it is actually marketed, and whether it is reimbursed, since these three statuses do not always align even for the medications approved for alcohol use disorder. This publication does not address any of this, and clinicians should verify the regulatory status of these molecules, and of the medications approved for alcohol use disorder, under the rules of their own jurisdiction.

Editorial collections

Tags

Verified on August 29, 2026 against the full text of the publication and its supplementary material where available. This analysis underwent an independent double reading. The English version was checked for conformity on September 21, 2026, against the figures of the French version and against the source. How we verify what we publish

This analysis is intended for healthcare professionals. It does not constitute a prescribing recommendation and does not replace individual clinical judgment.

Analysis from Psychiatry Evidence Base, evidence-based psychiatry, explained with rigor.

Report an error in this analysis

Follow Dr Stroescu on LinkedIn, for the review every Saturday