Rigour and Transparency

Last updated on 2026-07-10 | Edit this page

Overview

Questions

  • What is open source AI?
  • What are the limitation on explainability?
  • What do I need to declare about my use of Gen AI when publishing my work?

Objectives

  • Awareness of reproducibility concerns re: GenA I outputs
  • Awareness of limitations of explainable AI
  • Awareness of business models and how these impact behaviours
  • Awareness of differing journal requirements re: declaring use of Generative AI

What is transparency?


The UK Research Integrity Office define Transparency as a means of promoting trust and confidence.

This is demonstrated through:

  • reporting full methods,
  • publishing all results,
  • sharing data, code and materials,
  • and declaring conflicts of interest.

This includes acknowledging the use of tools such as emerging technologies, e.g. Generative AI.

“If you don’t pay for it you are the product”

Margaret McCartney, 2018

Open Acccess versus Open Source


There is a wide range of Generative AI and more specifically, Large Language Models that are “free” to use without registering for an account. These often offer very little control over what data is collected during use and re-used for future model training.

Although subscriptions provide a level of control over what data is collected about your use, their is limited information available about how the tool was created. They are often closed source.

How do you know you have used an appropriate method, if you do not know how the tool is deciding what to present to you?

Discussion

Think, Pair, Share

Read the definition of open source AI.

  • How does the Open Source AI definition compare to the Open Source definition?

  • How much does it reduce blackbox and aid explainability?

Spend 2 minutes thinking about your response.

Take it in turns to share your thoughts with your neighbours.

Explainable AI

Explainable AI refers to: the set of processes and methods that allows human users to comprehend and trust the results and output created by machine learning algorithms

Is Open Source AI explainable AI?

For more information about explainable AI see: What is Explainable AI? - Software Engineering Institute, Carnegie Mellon University

Rigour, disciplinary norms and the attention economy


The UK Research Integrity Office state that Rigour is demonstrated by behaviour that is in line with prevailing disciplinary norms and standards, including the use of appropriate methods.

“Attention is a resource — a person has only so much of it.”

Matthew Crawford, 2015

Many Generative AI tools are designed to keep the user engaged with the tool. You may have noticed one or more follow-up questions to your initial request.

Challenge

Challenge: Task review

The question I would like to answer: Is there a statistically significant relationship between population density, GDP per capita and life expectancy?

Watch the video of the interaction with an LLM: https://youtu.be/cecGWV6BpVo

This video does not have any audio

Alternatively, review the chat history: ChatGPT transcript

Answer the questions:

  • Was the question answered?
  • Where may the question answer of deviated?
  • Was there a risk of creating a non-reproducible result?

Generative AI tools may ask follow-up questions which could potentially lead you away from your data analysis aims.

Always ask a model how it has come to a conclusion and share the code it has used to generate the reportes repsonse.

This will enable you to verify the model outputs.

Publishing your research outputs


Research can take many forms. Traditional papers, software papers, software packages, contributions to open source projects or blog posts.

In May 2026, Nature reported on an increase of fake references in Bio-medical science papers. A study audited 2.5 million papers available via the PubMed Central (PMC) Open Access database published between January 2023 and February 2026. 300,000 were found to create fake references.

See: Surge in fake citations uncovered by audit of 2.5 million biomedical-science papers

Callout

Do falsified references constitute a breach of research integrity?

Do you think a paper should be retracted for including falsified references?

When submitting to journals or contributing to open source software projects, there will be guidance on how to contribute and increasingly, requirements to declare Generative AI use.

For example, Cambridge University Press require Author’s to declare and clearly explain their use of AI. Authors are also accountable for any use of AI.

The Journal for Open Source Software (JOSS), provides guidance on AI use for both Authors and Reviewers.

Do you agree with the different guidance for authors and reviewers?

Interestingly, the Comprehensive R Archive Network (CRAN) Repository Policy does not mention AI. However, “The ownership of copyright and intellectual property rights of all components of the package must be clear and unambiguous”, and authors are responsbible for ensuring there is no infringement or misrepresentation of copyright or licences.

Discussion

Check the requirements

Visit the website of a prominent journal in your field.

Does the author guidelines have any information about about AI generated content?

An example is the BERA Journal Author Guidelines.

Is there anything that surprises you?

What are the implications for your work?

Spend 2 minutes thinking about your response.

Take it in turns to share your thoughts with your neighbours.

Key Points
  • It is important to be transparent about where and how Generative AI has been used in your work flow
  • There are differences between the Open Source and Open AI definitions
  • Some Generative AI tools are designed to encourage engagement with the tool
  • Check the rules around Generative AI use for your choden dissemination channel