“We recently received a double-digit number of submissions from authors in the Far East… where we had strong evidence of machine generated content. Nothing made much sense in these papers, the authors, with their yahoo and aol type email addresses, frequently supplied false university affiliations, among other glaring problems. I suspected that these submissions were a test of our detection systems more so than anything else.”

That is Udo Schuklenk, professor of philosophy at Queen’s University, writing in his capacity as one of the editors of the journal, Bioethics.

His editorial was prompted by his journal’s publication of a critique of an article in the Journal of Medical Ethics (JME) for, among other things, multiple citations to “bibliographic fictions” (seemingly the result of unacknowledged AI use), which led to the latter article’s retraction.

The now-retracted JME article is “Panem, corticoids and circenses: the ethical fallout of Enhanced Games,” by Alexis Demas (who, Schuklenk notes, used a yahoo email address and whose stated institutional affiliation could not be verified).

The critique of the JME article, published (open access) in Bioethics, is by Ognjen Arandjelović (University of St. Andrews), and is titled “Against Moral Panic and Citation Fiction: A Critique of “Panem, Corticoids and Circenses” and a Proposal for Editorial Gatekeeping on Reference Integrity“.

Arandjelović discusses the reference failures and what to do about them in section 4 of his article, and I recommend people look at it. Here’s the introduction to that section:

It is tempting to treat reference failures as a mechanistic problem that could be addressed through technology and bureaucracy by installing more automated checks, tightening and enforcing checklists, reminding authors of their obligations. Not only banal, that response misdiagnoses the failure mode. In a publication culture where fast-turnaround commentaries are rewarded, reviewers often treat references as decorative rather than load-bearing, and journals rarely suffer meaningful consequences for bibliographic unreliability, low-grade fabrication, and high-grade sloppiness become rational strategies.

The core issue, then, is not attempting to help authors avoid mistakes. Authors know this. Rather, the challenge is in making it costly (in probability and consequence) to submit work with citations that do not exist or do not support what they are cited for. In other words, what is needed is deterrence, not assistance; accountability, not coaching.

Schuklenk describes how the case looks from the point of view of the editor:

Given the number of made-up references in the Journal of Medical Ethics paper, it is reasonable to assume that these references were the result of an AI hallucinating, a known phenomenon of AI-written content that adds non-existent references to the content the AI generated, and it is not unreasonable to suggest that the author of this manuscript was careless enough to leave a plethora of fake references in the manuscript. This doesn’t happen by accident. During the proof corrections stage of the article the author would have had ample time to correct any errors but chose not to. Your guess about how much else of the manuscript was generated using AI is as good as mine. We don’t currently have detection tools available to assist us in making reliable determinations. AI use also wasn’t disclosed by this author, if this is what is (very likely) at the heart of the fake references. I’m pleased to say that, apparently unlike the BMJ group of journals, Wiley, the publisher of this journal, has in place a highly sophisticated automated reference check that is available to the editorial team. This manuscript would not have gone out for peer review if it had been submitted to Bioethics, because it would have been eliminated after the reference (and possibly the AI generated content) screening. Surprisingly, the BMJ group of journals doesn’t seem to possess this sort of capacity, or it hasn’t been deployed in this instance.

I don’t know how the reader views the fact that the reviewers and editors of the Journal of Medical Ethics didn’t undertake a reference check for this manuscript given some of the already mentioned other ‘red flags’ that were in place. Arandelović thinks that they failed the readers of the journal, and ultimately that the journal failed as an institution on this occasion. It is easy to point fingers here, but truth be told, it has become exceedingly difficult to find willing reviewers for the ever-increasing number of papers that are prima facie worthy of peer review, and those reviewers are still expected to work pro bono, because publishers don’t wish to pay for their services. I am not surprised that reviewers do not spend their volunteer time undertaking detailed (or any) reference checks, especially given that the publisher could have systems in place that undertake that sort of task automatically.

Schuklenk’s editorial is here.

(via Johann Go)