ArXiv, a prominent platform for preprint research, is intensifying measures against the careless application of large language models in scientific publications.
Even though papers are shared on the site before undergoing peer review, arXiv (pronounced “archive”) has emerged as a vital conduit for research dissemination within disciplines such as computer science and mathematics. The platform has also become a valuable resource for analyzing scientific research trends.
To address the influx of substandard, AI-generated submissions, arXiv has introduced practices requiring first-time submitters to obtain endorsements from established researchers. Additionally, after more than two decades under Cornell’s management, the organization is transitioning to an independent nonprofit status. This change is expected to assist in increasing funding to tackle challenges like AI-generated issues.
In a recent update, Thomas Dietterich, chair of arXiv’s computer science section, announced via a post on Thursday that submissions lacking verification for the results generated by large language models (LLMs) compromise the integrity of the research. He emphasized, “if a submission contains incontrovertible evidence that the authors did not check the results of LLM generation, this means we can’t trust anything in the paper.”
Incontrovertible evidence might encompass instances of “hallucinated references” or exchanges involving the LLM, according to Dietterich. If such cases are identified, authors will face “a 1-year ban from arXiv, followed by the stipulation that future submissions must be accepted by a legitimate peer-reviewed venue.”
It’s important to clarify that this doesn’t equate to a blanket ban on LLM usage. Instead, as Dietterich mentioned, authors must assume “full responsibility” for their content, regardless of how it is produced. Consequently, researchers remain accountable if they directly incorporate “inappropriate language, plagiarized content, biased material, errors, mistakes, incorrect references, or misleading content” generated by an LLM.
Dietterich informed 404 Media that enforcement will follow a “one-strike” principle. However, moderators must flag any concerns, and section chairs will verify the evidence before any sanctions are levied. Authors will retain the right to appeal the ruling.
Recent studies have indicated a surge in fabricated citations within biomedical research, attributed to the influence of LLMs. This trend reveals that scientists aren’t alone in facing scrutiny; others have also been caught using AI-generated citations.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Why this news matters: ArXiv’s proactive stance in regulating the use of AI in research highlights the increasing concerns surrounding academic integrity. As reliance on AI technologies grows, ensuring the accuracy and accountability of scientific contributions becomes essential for the credibility of the research community.
#Research #repository #ArXiv #ban #authors #year #work