All in One View
Content from Introduction
Last updated on 2026-05-20 | Edit this page
Overview
Questions
- Why might a paper be retracted?
- How do paper retractions relate to integrity?
- What do we mean by research integrity?
Objectives
- Describe factors contributing to paper retractions
- Describe components of research integrity
- Explain how lack of integrity could lead to a paper retraction
Introducing Alma
Alma has recently joined a new research group as a junior member of the team. They are keen to make a good impression and be a valued contributor to the team. There has been a lot of buzz around the team and in general about AI, and how they may be helpful in the team’s current project. However, Alma isn’t so sure. Alma’s previous research team lost funding after it was found that the lead investigator had been falsifying data, and a number of papers were retracted. They are also aware of papers in high profile journals being found to have “hallucinated” citations, and concerns around wellbeing. Alma needs some help to critically evaluate if this will be an appropriate tool to use in their research process.
Throughout this lesson we will discuss a range of issues that will help Alma, and you, critically evaluate if an AI tool is appropriate to use in your projects.
Risks of retraction
In December 2023 Nature reported that more than 10,000 publications had been retracted during that calendar year alone. Publications can be retracted for a number of reasons. Some are retracted by the authors, others by journal editors. Some publications remain in the public conciousness long after they have been retracted.
An example of this is Wakefield’s 1998 article linking autism to MMR vaccines. This article was retracted by the Lancet in February 2010 after it was found that Dr Wakefield and colleagues were found to have acted unethically and that there were concerns about incorrect findings. It was also later discovered that the study was partially funded by lawyers acting for parents who were involved in lawsuits against vaccine manufacturers. The public health impact of this publication has been a reduction in vaccinations, leading to an increase in measles outbreaks, which can be life threatening.
For details see: Lancet retracts 12-year-old article linking autism to MMR vaccines
In addition, any articles that have been used in the training data sets for large language models will remain permanently embedded in the model. This can result in incorrect or misleading information being presented to users as facts via model output.
Retraction Watch
Retraction Watch is a blog that reports on the retraction of publications from journals and why these have occurred. In addition to the blog posts, there is also the database that stores information about which publications have been retracted, from where and when.
You may want to explore the blog and website, the leaderboards may be of particular interest.
What are the risks?
Which of the following could result in a paper being retracted?
- Hallucinated citations
- Conflict of interest
- Falsified data
- Misleading conclusions
All of these problems could/should require the paper to be retracted.
Research Integrity
Let’s consider these two definitions from the Oxford English Dictionary:
Research: Systematic investigation or inquiry aimed at contributing to knowledge of a theory, topic, etc., by careful consideration, observation, or study of a subject.
(https://doi.org/10.1093/OED/1194777451)
Integrity: Soundness of moral principle; the character of uncorrupted virtue, esp. in relation to truth and fair dealing; uprightness, honesty, sincerity.
(https://doi.org/10.1093/OED/1327125083)
What do these mean with regards to how we undertake research?
Tell Your neighbour: What does research integrity mean to you?
- Spend 2 minutes thinking about what research integrity means to you
- Share your understanding of research integrity with your neighbour
- Was anything different, or unexpected in their understanding?
The UK Research Integrity Office defines research integrity as all of the factors that underpin good research practice and promote trust and confidence in the research process. These are across the entire research process.
Dimensions of research integrity

Honesty
In all aspects of research, including:
- Planning
- Methods
- Data collection
- Credit
- Reporting
- Interpretation
Transparency
Promoting trust and confidence, including by:
- Reporting full methods
- Publishing all results
- Sharing data, code and materials
- Declaring conflicts of interest
Accountability
Of everyone involved in research, including:
- Researchers
- Institutions
- Funding bodies
- Publishers
These five principles sit alongside the four principles defined in The Singapore Statement agreed at the 2010 World Conference on Research Integrity (WCRI):
- Honesty in all aspects of research
- Accountability in the conduct of research
- Professional courtesy and fairness in working with others
- Good stewardship of research on behalf of others
We will explore each of these dimensions with regards to the use of Generative AI.
Integrity and retractions
If we know reconsider Wakefield (1998), there were a number of concerns regarding research integerity.
What were the main concerns?
Do these overlap with retraction scenarios?
Minute Paper: Integrity and retractions
Write for 1 minute on the topic of how a lack of integrity could lead to a paper retraction.
- Papers may be retracted by an author or the journal’s editorial team.
- Integrity concerns may be a factor in retraction.
- There are multiple dimensions to integrity.
Content from Rigour and Transparency
Last updated on 2026-07-10 | Edit this page
Overview
Questions
- What is open source AI?
- What are the limitation on explainability?
- What do I need to declare about my use of Gen AI when publishing my work?
Objectives
- Awareness of reproducibility concerns re: GenA I outputs
- Awareness of limitations of explainable AI
- Awareness of business models and how these impact behaviours
- Awareness of differing journal requirements re: declaring use of Generative AI
What is transparency?
The UK Research Integrity Office define Transparency as a means of promoting trust and confidence.
This is demonstrated through:
- reporting full methods,
- publishing all results,
- sharing data, code and materials,
- and declaring conflicts of interest.
This includes acknowledging the use of tools such as emerging technologies, e.g. Generative AI.
“If you don’t pay for it you are the product”
Margaret McCartney, 2018
Open Acccess versus Open Source
There is a wide range of Generative AI and more specifically, Large Language Models that are “free” to use without registering for an account. These often offer very little control over what data is collected during use and re-used for future model training.
Although subscriptions provide a level of control over what data is collected about your use, their is limited information available about how the tool was created. They are often closed source.
How do you know you have used an appropriate method, if you do not know how the tool is deciding what to present to you?
Explainable AI
Explainable AI refers to: the set of processes and methods that allows human users to comprehend and trust the results and output created by machine learning algorithms
Is Open Source AI explainable AI?
For more information about explainable AI see: What is Explainable AI? - Software Engineering Institute, Carnegie Mellon University
Rigour, disciplinary norms and the attention economy
The UK Research Integrity Office state that Rigour is demonstrated by behaviour that is in line with prevailing disciplinary norms and standards, including the use of appropriate methods.
“Attention is a resource — a person has only so much of it.”
Many Generative AI tools are designed to keep the user engaged with the tool. You may have noticed one or more follow-up questions to your initial request.
Challenge: Task review
The question I would like to answer: Is there a statistically significant relationship between population density, GDP per capita and life expectancy?
Watch the video of the interaction with an LLM: https://youtu.be/cecGWV6BpVo
This video does not have any audio
Alternatively, review the chat history: ChatGPT transcript
Answer the questions:
- Was the question answered?
- Where may the question answer of deviated?
- Was there a risk of creating a non-reproducible result?
Generative AI tools may ask follow-up questions which could potentially lead you away from your data analysis aims.
Always ask a model how it has come to a conclusion and share the code it has used to generate the reportes repsonse.
This will enable you to verify the model outputs.
Publishing your research outputs
Research can take many forms. Traditional papers, software papers, software packages, contributions to open source projects or blog posts.
In May 2026, Nature reported on an increase of fake references in Bio-medical science papers. A study audited 2.5 million papers available via the PubMed Central (PMC) Open Access database published between January 2023 and February 2026. 300,000 were found to create fake references.
See: Surge in fake citations uncovered by audit of 2.5 million biomedical-science papers
Do falsified references constitute a breach of research integrity?
Do you think a paper should be retracted for including falsified references?
When submitting to journals or contributing to open source software projects, there will be guidance on how to contribute and increasingly, requirements to declare Generative AI use.
For example, Cambridge University Press require Author’s to declare and clearly explain their use of AI. Authors are also accountable for any use of AI.
The Journal for Open Source Software (JOSS), provides guidance on AI use for both Authors and Reviewers.
Do you agree with the different guidance for authors and reviewers?
Interestingly, the Comprehensive R Archive Network (CRAN) Repository Policy does not mention AI. However, “The ownership of copyright and intellectual property rights of all components of the package must be clear and unambiguous”, and authors are responsbible for ensuring there is no infringement or misrepresentation of copyright or licences.
Check the requirements
Visit the website of a prominent journal in your field.
Does the author guidelines have any information about about AI generated content?
An example is the BERA Journal Author Guidelines.
Is there anything that surprises you?
What are the implications for your work?
Spend 2 minutes thinking about your response.
Take it in turns to share your thoughts with your neighbours.
- It is important to be transparent about where and how Generative AI has been used in your work flow
- There are differences between the Open Source and Open AI definitions
- Some Generative AI tools are designed to encourage engagement with the tool
- Check the rules around Generative AI use for your choden dissemination channel
Content from Implications for Learning
Last updated on 2026-07-10 | Edit this page
Overview
Questions
- What is a shallow approach to learning?
- What is a deep approach to learning?
- How does the use of GenAI for coding impact learning?
- How can GenAI be used as a tool to support critical coding rather than a shortcut?
Objectives
- Understand the difference between a shallow approach and a deep approach to learning
- Understand how genAI can enhance or hinder skill acquisition
- Relate the implications for using GenAI to the act of learning how to code
Introduction
In this episode we will take a closer look at different approaches to learning and how they relate to using GenAI in research software development.
The Shallow vs Deep Approach to Learning
Consider the following questions:
- How many legs does a spider have?
- How is a spider not an insect?
Shallow learning, also sometimes called surface learning or rote learning, refers to learning activities that are characterized by recalling and rote memorization. The knowledge gained from shallow learning is considered passive and tends to fade away from our memory.
Deep learning refers to learning activities that are characterized by an effort to connect with and understand the material conceptually. The knowledge gained from deep learning is considered active. Analyzing meaning by drawing connections, elaborating on ideas, and linking them to prior knowledge, which leads to long-term retention in memory.
Importantly, slow speed is an inherent characteristic of deep thinking and learning.
Up-skill and De-skill
The notion that thinking cannot be sped up stands in stark contrast to the promise of GenAI to accelerate productivity and make things faster. Having GenAI perform a task for you is often described as cognitive offloading.
In a 2026 study, Anthropic sought to determine whether cognitive offloading can prevent people from growing their coding skills. In a randomized controlled trial the study found that using AI assistance led to a statistically significant decrease in mastery.
A deep learning approach requires you to be aware of how you are thinking and learning. The learning process becomes a meaningful experience in and of itself.
In June 2026 the Norwegian government announced plans to reduce access to AI tools by learners in schools.
From August 2026 these tools would not be available for learners aged 6 to 13yrs, with those aged 14 and 16 only able to use it under direct teacher supervision.
Learners aged 17 to 19yrs will learn to “use AI appropriately” to prepare them for further education and work.
The Norwegian Prime Minister Jonas Gahr Støre said that using AI increases the risk that young children miss important steps in their education:“The most important thing in school is that our children learn to read, write, and do mathematics,”.
AI Engagement Strategies
Over-reliance on AI for coding can prevent researchers from developing essential skills in research software development and data analysis. Without a solid understanding of the code it is impossible to reliably verify whether research results are correct and valid.
The study by Anthropic mentioned above also found that how someone used AI had a great impact on how much information they retained. Coders who used AI assistance not just to produce code but to build comprehension while doing so showed stronger mastery.
Challenge:
What are some strategies to avoid cognitive offloading and instead use AI to strengthen your research computing skills?
Here are some strategies that could help build comprehension.
- Ask follow-up questions
- Request explanations
- Ask for code generation along with explanations of the generated code
- Pose conceptual questions then use improved understanding to complete the task
- Ask for code review
- Use “what if” and “how/why” questions
- Introduce “friction”: Instruct AI to push back, to never provide an answer outright, etc
- Take it slow!
At the Association for Learning Development Conference 2025, a workshop explored how activities could be designed to enable learners to engage critically with AI.
Read the workshop write-up:What is lazy metacognition and what can we do about it?
Your overall goal should always be to focus on your learning. That means to adopt practices that help you keep learning to code and to be able to scrutinize the answers given back by GenAI.
GenAI output often sounds very confident, which can make us inclined to accept it as correct and factual without critically evaluating it. However, it is important to treat AI outputs as suggestions rather than solutions. Also, remember that you as the researcher need to take responsibility for any AI-generated code you use. We will come back to this in the next episode.
- Shallow learning reduces attention span and hampers deep learning that requires focus and critical thinking.
- Deep learning is focused on problem-solving and connecting various sources of new and existing knowledge.
- GenAI carries the danger of the user being content with shallow learning.
- Avoid shallow learning by using GenAI mindfully.
Content from Respect and Accountability
Last updated on 2026-02-05 | Edit this page
Overview
Questions
- How do you write a lesson using Markdown and sandpaper?
Objectives
- Describe key environmental concerns relating to GenAI and machine learning
- State key issues relating to data workers and workers rights
- Awareness of embedded values in the tools
- Describe current legal and ethical debates relating to data acquisition for training
Introduction
This is a lesson created via The Carpentries Workbench. It is written in Pandoc-flavored Markdown for static files and R Markdown for dynamic files that can render code into output. Please refer to the Introduction to The Carpentries Workbench for full documentation.
What you need to know is that there are three sections required for a valid Carpentries lesson:
-
questionsare displayed at the beginning of the episode to prime the learner for the content. -
objectivesare the learning objectives for an episode displayed with the questions. -
keypointsare displayed at the end of the episode to reinforce the objectives.
Challenge 1: Can you do it?
What is the output of this command?
R
paste("This", "new", "lesson", "looks", "good")
OUTPUT
[1] "This new lesson looks good"
Challenge 2: how do you nest solutions within challenge blocks?
You can add a line with at least three colons and a
solution tag.
Figures
You can use standard markdown for static figures with the following syntax:
{alt='alt text for accessibility purposes'}
Callout sections can highlight information.
They are sometimes used to emphasise particularly important points but are also used in some lessons to present “asides”: content that is not central to the narrative of the lesson, e.g. by providing the answer to a commonly-asked question.
Math
One of our episodes contains \(\LaTeX\) equations when describing how to create dynamic reports with {knitr}, so we now use mathjax to describe this:
$\alpha = \dfrac{1}{(1 - \beta)^2}$ becomes: \(\alpha = \dfrac{1}{(1 - \beta)^2}\)
Cool, right?
- Use
.mdfiles for episodes when you want static content - Use
.Rmdfiles for episodes when you need to generate output - Run
sandpaper::check_lesson()to identify any issues with your lesson - Run
sandpaper::build_lesson()to preview your lesson locally