Earlier this year, investigators and research compliance officials grappling with the brave, new (and increasingly frightening) world of AI received some assistance when NIH made a blog post and the HHS Office of Research Integrity (ORI) released guidance on the use of generative AI as it relates to misconduct.[i] The guidance will be followed by “best practices to prevent AI-related research misconduct and guidance on the appropriate use of AI tools throughout the research process,” ORI promised.
There’s no doubt that AI use poses challenges, but one type is particularly low-hanging fruit—hallucinated citations or references. Yet, these don’t always constitute misconduct. “When references do not serve as data or results, fabricated citations may not constitute fabricated data or results,” according to ORI’s guidance. “However, in some cases, such as literature reviews, references are data because they arise from scientific inquiry. In these cases, fabricated citations could be the subject of research misconduct allegations.”
NIH didn’t precisely land on the “yes if data” argument, saying “presenting AI-generated, non-existent references overtly as real…could constitute data fabrication.” However, it linked to a paper published in April by David B. Resnik and Mohammad Hosseini, which makes exactly that point.[ii] Resnik is a bioethicist with the National Institute of Environmental Health Sciences.
In an exclusive interview after ORI’s guidance was issued, Hosseini, an assistant professor at Northwestern University Feinberg School of Medicine’s Department of Preventive Medicine, told RRC he is “really glad the topic of hallucinated citations has received a lot of attention.” He also discussed what he likes about ORI and NIH’s guidance, how the use of AI detection software may have the unintended consequence of making fabrications harder to catch and what Northwestern has implemented that puts it in “a good position” to oversee AI use.
RRC: First, how did you become interested in studying AI in general and hallucinations of citations, specifically?
Hosseini: I was working on emerging technologies even before ChatGPT was released, and as someone who has worked for more than a decade on authorship issues, the first natural question I got interested in was the issue of AI authorship. I have also been interested in and worked on citation ethics and inaccurate citations. So, when ChatGPT and other AI systems became mainstream, a reasonable question for me was about the ethics of their use in citation practices. I first wrote about hallucinated citations in 2024.
RRC: How did you feel when you saw ORI’s guidance? Were you involved in its development?
Hosseini: My immediate reaction was: “Yeah, finally.” No, I was not.
For the last three years, we’ve been teaching responsible conduct of research for the National Institute of Allergy and Infectious Diseases. This is mandatory for everybody. We have been teaching them using case studies and so on about how to use AI in a responsible way. Although the NIH had guidance on its use, I think [ORI’s] is more specific and it addresses some of the specific issues on misconduct. That being said, I have my own opinions.
RRC: About what?
Hosseini: I was a little bit surprised to see a few things. JAMA’s most recent guidance on the use of AI [published Aug. 10] said if you use AI for grammar and spelling correction, basically for copy editing, you don’t need to disclose it.[iii] But the ORI guidance says researchers should disclose any generative AI use…so, to me, that reads that you must disclose your use of AI even if it’s only for grammar and copy editing, which is a little demanding. ORI, in my understanding, always errs on the side of caution and transparency, given the significance and sensitivity of the topics and the research they deal with. They want to make sure that all kinds of uses are basically disclosed and people are being as open and transparent as possible.
One thing that is very important to notice is that these guidances come and go. Because the landscape is evolving so quickly, guidance also needs to change and evolve. But I think [the ORI guidance] is really good. It covers quite a lot. I like that it provides guidance not just for researchers, but also for those who are investigating allegations of misconduct.
RRC: Can you provide an example?
It says, “When evaluating allegations involving generative AI, institutional committees should gather and consider all available evidence. Institutional investigation committees must include persons with appropriate scientific expertise in the respondent’s research field.” This may sound subtle, but it’s very important because it highlights the fact that not just any investigator can do any specific [misconduct] investigation. You have to have the expertise [in] the field and know how generative AI can be used and the risks associated with that specific use case.
It also says that “institutional committee members with appropriate scientific expertise in the respondent’s research field can help establish whether a particular use of a gen AI tool falls within the accepted practices of the relevant research community. Institutions could consider including [a] subject matter expert in AI when creating inquiry and/or investigation committees to provide additional insights into the uses and functionality of generative AI.” I like this.
I was glad to see that they specifically mentioned that the output of generative AI can be evidence of research misconduct [and] highlight that plagiarism detection tools have limitations. In [a recently published paper], we talk about idea plagiarism. These tools, like plagiarism detection tools, can spot verbatim plagiarism. But with generative AI, you can paraphrase and rephrase forever until signs of verbatim plagiarism are completely gone.
It is really cool that the guidance is actually highlighting that as a limitation because a lot of publishers keep hyping their plagiarism detection software tool and saying, “Oh, we use this-and-that tool.” But in fact, those tools at this moment, in my understanding, are not really helpful because people can keep paraphrasing. And in many institutions, people have access to plagiarism detection software. So, if someone wants to copy-paste something from a paper that has been published in the past and asks generative AI to rephrase it, they can then upload the new content to the plagiarism detection software that is provided for free by the university to see whether there [are] any signs of plagiarism.
From a university point of view, I think this is good practice because it prevents allegations of misconduct. It prevents reputational damage and so on. But if you think about it, it is almost like a backdoor that facilitates further misuse of previous publications or content that has been published previously.
RRC: Are there other parts of the guidance you liked?
Hosseini: The bit about the burden of proof was also interesting. The guidance says that the burden of proving honest error under this regulation is with the researcher. The implication is that researchers should save all the prompts and outputs and whatever record they have checked, because, if ever an allegation is landed and an investigation has started, the difference between someone who would be accused of honest error versus someone who’s accused of misconduct is the trail, is the record.
Actually, I was teaching today, and based on this, I was telling my students that you have to save everything. Don’t delete stuff because, if it comes to an investigation of misconduct, the logs you have with the generative AI you have used could be the difference between you getting cleared of an allegation of misconduct or you being accused of misconduct, which has serious consequences.
RRC: On LinkedIn, some people were commenting that they thought there should have been stronger condemnation of hallucinated citations. And you were explaining that they’re unethical, but they may not be true misconduct except in two instances, two types of studies.
Hosseini: It seemed the person, or some of the people who are just shouting so loudly on LinkedIn, had just not read the paper. We specifically say that we are dealing with misconduct in a legal sense. And misconduct in a legal sense has a specific definition. I cannot just haphazardly apply stuff and say, “Oh, based on what I feel, this should be misconduct.” The fact that our paper has attracted attention and started a debate or has brought the issue to the radar of other scholars is great. And if other people disagree with it, by all means, share [those views].
RRC: Is this emphasis on the citations missing the mark at all? Is it really the most worrisome possible misuse of AI in research?
Hosseini: Misuses can be different in many different contexts. I don’t think anyone can come up with a sort of reliable spectrum that offers judgment calls like “This is worse than this; this misuse is worse than that, or inaccurate citations are better than this.” I don’t think these rankings are helpful either. At the same time, hallucinated citations have been very visible.
I’m really glad the topic of hallucinated citations has received a lot of attention. One thing that makes it very different from many other use cases is that it is easy to detect. People can just simply run a few lines of code to compare the references of a paper with already existing bibliometric indexes to see what reference exists and what reference doesn’t. But again, how do you make a distinction between typos or [a] simple, honest, inaccurate reference that has a bibliometric error? Say the year is wrong, or the name of the author is wrong. How do you make a distinction between that and a hallucinated citation that is generated by an AI system?
RRC: There was a widely criticized paper by the Make America Healthy Again Commission that was found to have numerous hallucinated citations. But it was not clear these could have been termed research misconduct.
Hosseini: People don’t realize that misconduct is a serious allegation, is a serious matter. If you are [accused], if there’s overwhelming evidence that you have committed misconduct, there’s a whole disciplinary process that needs to kick in. And if every single fabricated citation is an instance of misconduct, we need thousands of—and I’m not joking—we need thousands of people to just follow on with those cases and enforce existing regulation. If we don’t do that, then it feels like everything that has been tagged as misconduct is a joke, because it suggests that we can tag something as misconduct, and there will be no consequences because we simply don’t have the resources to follow through.
RRC: Did you run this concept of when a hallucinated citation is misconduct by research integrity officers?
Hosseini: We were going back and forth with this paper for months. We shared it with so many people from different areas, people who do different things within universities. If you discuss this with, for example, research integrity officers within universities who have to enforce existing misconduct regulations, this, for them, even at the level of only naming hallucinated citations in literature reviews as an instance of misconduct, even this is a logistical nightmare because they don’t have the capacity to follow through. And if you call it misconduct and don’t follow through, that is much worse than not calling it misconduct.
In the paper, we suggested that [sanctions] should be up to the institutions: “We believe that it should be up to the research institutions to devise appropriate sanctions for researchers who include GenAI-generated hallucinated citations in research documents. In supporting this view, we believe that publishers and journals should invest in automated tools that detect hallucinated citations, and that they should report them to researchers’ institution[s], if appropriate. Potential penalties should be proportionate to the nature, degree and frequency of a researcher’s offense.”
RRC: In your experience, do a lot of misconduct allegations actually implicate AI, citation-wise or otherwise, at the moment?
Hosseini: I think you would have to discuss that with people who are doing misconduct investigations.
RRC: Can you talk about how Northwestern approaches the use of AI?
Hosseini: We have a good foundation for AI ethics. I am the director of the AI Ethics Committee for the Northwestern Network of Collaborative Intelligence, engaging faculty and staff from both campuses and across fields and departments. AI ethics is taken seriously. Our library has AI literacy programs. The Office of Provost has AI literacy programs for teachers. There’s guidance for students, researchers, teachers.
So, there’s three different kinds of guidance depending on what you do and where you are within the university. iThenticate plagiarism detection software is provided for free to everybody, as is Proofig AI, which is for detecting image manipulation.
We all have access to Microsoft Copilot. Compared with many other institutions that don’t even have guidance, we are in a good position, and the university is trying to stay up to date and adjust itself within this evolving landscape. ✧
