All in One View

Content from Introduction


Last updated on 2026-05-20 | Edit this page

Overview

Questions

  • Why might a paper be retracted?
  • How do paper retractions relate to integrity?
  • What do we mean by research integrity?

Objectives

  • Describe factors contributing to paper retractions
  • Describe components of research integrity
  • Explain how lack of integrity could lead to a paper retraction

Introducing Alma


Alma has recently joined a new research group as a junior member of the team. They are keen to make a good impression and be a valued contributor to the team. There has been a lot of buzz around the team and in general about AI, and how they may be helpful in the team’s current project. However, Alma isn’t so sure. Alma’s previous research team lost funding after it was found that the lead investigator had been falsifying data, and a number of papers were retracted. They are also aware of papers in high profile journals being found to have “hallucinated” citations, and concerns around wellbeing. Alma needs some help to critically evaluate if this will be an appropriate tool to use in their research process.

Throughout this lesson we will discuss a range of issues that will help Alma, and you, critically evaluate if an AI tool is appropriate to use in your projects.

Risks of retraction


In December 2023 Nature reported that more than 10,000 publications had been retracted during that calendar year alone. Publications can be retracted for a number of reasons. Some are retracted by the authors, others by journal editors. Some publications remain in the public conciousness long after they have been retracted.

An example of this is Wakefield’s 1998 article linking autism to MMR vaccines. This article was retracted by the Lancet in February 2010 after it was found that Dr Wakefield and colleagues were found to have acted unethically and that there were concerns about incorrect findings. It was also later discovered that the study was partially funded by lawyers acting for parents who were involved in lawsuits against vaccine manufacturers. The public health impact of this publication has been a reduction in vaccinations, leading to an increase in measles outbreaks, which can be life threatening.

For details see: Lancet retracts 12-year-old article linking autism to MMR vaccines

In addition, any articles that have been used in the training data sets for large language models will remain permanently embedded in the model. This can result in incorrect or misleading information being presented to users as facts via model output.

Callout

Retraction Watch

Retraction Watch is a blog that reports on the retraction of publications from journals and why these have occurred. In addition to the blog posts, there is also the database that stores information about which publications have been retracted, from where and when.

You may want to explore the blog and website, the leaderboards may be of particular interest.

Challenge

What are the risks?

Which of the following could result in a paper being retracted?

  1. Hallucinated citations
  2. Conflict of interest
  3. Falsified data
  4. Misleading conclusions

All of these problems could/should require the paper to be retracted.

Research Integrity


Let’s consider these two definitions from the Oxford English Dictionary:

Research: Systematic investigation or inquiry aimed at contributing to knowledge of a theory, topic, etc., by careful consideration, observation, or study of a subject.

(https://doi.org/10.1093/OED/1194777451)

Integrity: Soundness of moral principle; the character of uncorrupted virtue, esp. in relation to truth and fair dealing; uprightness, honesty, sincerity.

(https://doi.org/10.1093/OED/1327125083)

What do these mean with regards to how we undertake research?

Discussion

Tell Your neighbour: What does research integrity mean to you?

  1. Spend 2 minutes thinking about what research integrity means to you
  2. Share your understanding of research integrity with your neighbour
  3. Was anything different, or unexpected in their understanding?

The UK Research Integrity Office defines research integrity as all of the factors that underpin good research practice and promote trust and confidence in the research process. These are across the entire research process.

Dimensions of research integrity

Dimensions of Research Integrity
Dimensions of Research Integrity
Honesty

In all aspects of research, including:

  • Planning
  • Methods
  • Data collection
  • Credit
  • Reporting
  • Interpretation
Transparency

Promoting trust and confidence, including by:

  • Reporting full methods
  • Publishing all results
  • Sharing data, code and materials
  • Declaring conflicts of interest
Accountability

Of everyone involved in research, including:

  • Researchers
  • Institutions
  • Funding bodies
  • Publishers
Respect

For everyone & everything involved in research, including:

  • Colleagues
  • Other researchers
  • Participants
  • Animals
  • The environment
Rigour

In line with disciplinary norms, including in:

  • Appropriate methods
  • Following protocols
  • Interpreting data
  • Drawing conclusions
  • Disseminating

These five principles sit alongside the four principles defined in The Singapore Statement agreed at the 2010 World Conference on Research Integrity (WCRI):

  • Honesty in all aspects of research
  • Accountability in the conduct of research
  • Professional courtesy and fairness in working with others
  • Good stewardship of research on behalf of others

We will explore each of these dimensions with regards to the use of Generative AI.

Integrity and retractions


If we know reconsider Wakefield (1998), there were a number of concerns regarding research integerity.

What were the main concerns?

Do these overlap with retraction scenarios?

Discussion

Minute Paper: Integrity and retractions

Write for 1 minute on the topic of how a lack of integrity could lead to a paper retraction.

Key Points
  • Papers may be retracted by an author or the journal’s editorial team.
  • Integrity concerns may be a factor in retraction.
  • There are multiple dimensions to integrity.

Content from Rigour and Transparency


Last updated on 2026-07-10 | Edit this page

Overview

Questions

  • What is open source AI?
  • What are the limitation on explainability?
  • What do I need to declare about my use of Gen AI when publishing my work?

Objectives

  • Awareness of reproducibility concerns re: GenA I outputs
  • Awareness of limitations of explainable AI
  • Awareness of business models and how these impact behaviours
  • Awareness of differing journal requirements re: declaring use of Generative AI

What is transparency?


The UK Research Integrity Office define Transparency as a means of promoting trust and confidence.

This is demonstrated through:

  • reporting full methods,
  • publishing all results,
  • sharing data, code and materials,
  • and declaring conflicts of interest.

This includes acknowledging the use of tools such as emerging technologies, e.g. Generative AI.

“If you don’t pay for it you are the product”

Margaret McCartney, 2018

Open Acccess versus Open Source


There is a wide range of Generative AI and more specifically, Large Language Models that are “free” to use without registering for an account. These often offer very little control over what data is collected during use and re-used for future model training.

Although subscriptions provide a level of control over what data is collected about your use, their is limited information available about how the tool was created. They are often closed source.

How do you know you have used an appropriate method, if you do not know how the tool is deciding what to present to you?

Discussion

Think, Pair, Share

Read the definition of open source AI.

  • How does the Open Source AI definition compare to the Open Source definition?

  • How much does it reduce blackbox and aid explainability?

Spend 2 minutes thinking about your response.

Take it in turns to share your thoughts with your neighbours.

Explainable AI

Explainable AI refers to: the set of processes and methods that allows human users to comprehend and trust the results and output created by machine learning algorithms

Is Open Source AI explainable AI?

For more information about explainable AI see: What is Explainable AI? - Software Engineering Institute, Carnegie Mellon University

Rigour, disciplinary norms and the attention economy


The UK Research Integrity Office state that Rigour is demonstrated by behaviour that is in line with prevailing disciplinary norms and standards, including the use of appropriate methods.

“Attention is a resource — a person has only so much of it.”

Matthew Crawford, 2015

Many Generative AI tools are designed to keep the user engaged with the tool. You may have noticed one or more follow-up questions to your initial request.

Challenge

Challenge: Task review

The question I would like to answer: Is there a statistically significant relationship between population density, GDP per capita and life expectancy?

Watch the video of the interaction with an LLM: https://youtu.be/cecGWV6BpVo

This video does not have any audio

Alternatively, review the chat history: ChatGPT transcript

Answer the questions:

  • Was the question answered?
  • Where may the question answer of deviated?
  • Was there a risk of creating a non-reproducible result?

Generative AI tools may ask follow-up questions which could potentially lead you away from your data analysis aims.

Always ask a model how it has come to a conclusion and share the code it has used to generate the reportes repsonse.

This will enable you to verify the model outputs.

Publishing your research outputs


Research can take many forms. Traditional papers, software papers, software packages, contributions to open source projects or blog posts.

In May 2026, Nature reported on an increase of fake references in Bio-medical science papers. A study audited 2.5 million papers available via the PubMed Central (PMC) Open Access database published between January 2023 and February 2026. 300,000 were found to create fake references.

See: Surge in fake citations uncovered by audit of 2.5 million biomedical-science papers

Callout

Do falsified references constitute a breach of research integrity?

Do you think a paper should be retracted for including falsified references?

When submitting to journals or contributing to open source software projects, there will be guidance on how to contribute and increasingly, requirements to declare Generative AI use.

For example, Cambridge University Press require Author’s to declare and clearly explain their use of AI. Authors are also accountable for any use of AI.

The Journal for Open Source Software (JOSS), provides guidance on AI use for both Authors and Reviewers.

Do you agree with the different guidance for authors and reviewers?

Interestingly, the Comprehensive R Archive Network (CRAN) Repository Policy does not mention AI. However, “The ownership of copyright and intellectual property rights of all components of the package must be clear and unambiguous”, and authors are responsbible for ensuring there is no infringement or misrepresentation of copyright or licences.

Discussion

Check the requirements

Visit the website of a prominent journal in your field.

Does the author guidelines have any information about about AI generated content?

An example is the BERA Journal Author Guidelines.

Is there anything that surprises you?

What are the implications for your work?

Spend 2 minutes thinking about your response.

Take it in turns to share your thoughts with your neighbours.

Key Points
  • It is important to be transparent about where and how Generative AI has been used in your work flow
  • There are differences between the Open Source and Open AI definitions
  • Some Generative AI tools are designed to encourage engagement with the tool
  • Check the rules around Generative AI use for your choden dissemination channel

Content from Implications for Learning


Last updated on 2026-07-10 | Edit this page

Overview

Questions

  • What is a shallow approach to learning?
  • What is a deep approach to learning?
  • How does the use of GenAI for coding impact learning?
  • How can GenAI be used as a tool to support critical coding rather than a shortcut?

Objectives

  • Understand the difference between a shallow approach and a deep approach to learning
  • Understand how genAI can enhance or hinder skill acquisition
  • Relate the implications for using GenAI to the act of learning how to code

Introduction


In this episode we will take a closer look at different approaches to learning and how they relate to using GenAI in research software development.

The Shallow vs Deep Approach to Learning


Consider the following questions:

  1. How many legs does a spider have?
  2. How is a spider not an insect?
Discussion

Turn to your neighbour and share your thoughts:

  1. How are these two questions different?
  2. How would you go about providing answers to these questions?
  3. What kind of learning takes place in each of these questions?
  4. What do you think you will remember about the topics tomorrow? Next month?

Shallow learning, also sometimes called surface learning or rote learning, refers to learning activities that are characterized by recalling and rote memorization. The knowledge gained from shallow learning is considered passive and tends to fade away from our memory.

Deep learning refers to learning activities that are characterized by an effort to connect with and understand the material conceptually. The knowledge gained from deep learning is considered active. Analyzing meaning by drawing connections, elaborating on ideas, and linking them to prior knowledge, which leads to long-term retention in memory.

Importantly, slow speed is an inherent characteristic of deep thinking and learning.

Up-skill and De-skill


The notion that thinking cannot be sped up stands in stark contrast to the promise of GenAI to accelerate productivity and make things faster. Having GenAI perform a task for you is often described as cognitive offloading.

In a 2026 study, Anthropic sought to determine whether cognitive offloading can prevent people from growing their coding skills. In a randomized controlled trial the study found that using AI assistance led to a statistically significant decrease in mastery.

A deep learning approach requires you to be aware of how you are thinking and learning. The learning process becomes a meaningful experience in and of itself.

Discussion

Turn to your neighbour and share your thoughts:

  1. When you look back at a recent interaction with an LLM, what parts of the thinking did you do, and what parts did you subconsciously let the AI do?
  2. How can you develop your ability to recognize your choices in the moment?
Callout

In June 2026 the Norwegian government announced plans to reduce access to AI tools by learners in schools.

From August 2026 these tools would not be available for learners aged 6 to 13yrs, with those aged 14 and 16 only able to use it under direct teacher supervision.

Learners aged 17 to 19yrs will learn to “use AI appropriately” to prepare them for further education and work.

The Norwegian Prime Minister Jonas Gahr Støre said that using AI increases the risk that young children miss important steps in their education:“The most important thing in school is that our children learn to read, write, and do mathematics,”.

Source: PC Mag UK: Norway bans AI in elementary schools

AI Engagement Strategies


Over-reliance on AI for coding can prevent researchers from developing essential skills in research software development and data analysis. Without a solid understanding of the code it is impossible to reliably verify whether research results are correct and valid.

The study by Anthropic mentioned above also found that how someone used AI had a great impact on how much information they retained. Coders who used AI assistance not just to produce code but to build comprehension while doing so showed stronger mastery.

Challenge

Challenge:

What are some strategies to avoid cognitive offloading and instead use AI to strengthen your research computing skills?

Here are some strategies that could help build comprehension.

  • Ask follow-up questions
  • Request explanations
  • Ask for code generation along with explanations of the generated code
  • Pose conceptual questions then use improved understanding to complete the task
  • Ask for code review
  • Use “what if” and “how/why” questions
  • Introduce “friction”: Instruct AI to push back, to never provide an answer outright, etc
  • Take it slow!
Callout

At the Association for Learning Development Conference 2025, a workshop explored how activities could be designed to enable learners to engage critically with AI.

Read the workshop write-up:What is lazy metacognition and what can we do about it?

Your overall goal should always be to focus on your learning. That means to adopt practices that help you keep learning to code and to be able to scrutinize the answers given back by GenAI.

GenAI output often sounds very confident, which can make us inclined to accept it as correct and factual without critically evaluating it. However, it is important to treat AI outputs as suggestions rather than solutions. Also, remember that you as the researcher need to take responsibility for any AI-generated code you use. We will come back to this in the next episode.

Key Points
  • Shallow learning reduces attention span and hampers deep learning that requires focus and critical thinking.
  • Deep learning is focused on problem-solving and connecting various sources of new and existing knowledge.
  • GenAI carries the danger of the user being content with shallow learning.
  • Avoid shallow learning by using GenAI mindfully.

Content from Respect and Accountability


Last updated on 2026-02-05 | Edit this page

Overview

Questions

  • How do you write a lesson using Markdown and sandpaper?

Objectives

  • Describe key environmental concerns relating to GenAI and machine learning
  • State key issues relating to data workers and workers rights
  • Awareness of embedded values in the tools
  • Describe current legal and ethical debates relating to data acquisition for training

Introduction


This is a lesson created via The Carpentries Workbench. It is written in Pandoc-flavored Markdown for static files and R Markdown for dynamic files that can render code into output. Please refer to the Introduction to The Carpentries Workbench for full documentation.

What you need to know is that there are three sections required for a valid Carpentries lesson:

  1. questions are displayed at the beginning of the episode to prime the learner for the content.
  2. objectives are the learning objectives for an episode displayed with the questions.
  3. keypoints are displayed at the end of the episode to reinforce the objectives.
Challenge

Challenge 1: Can you do it?

What is the output of this command?

R

paste("This", "new", "lesson", "looks", "good")

OUTPUT

[1] "This new lesson looks good"
Challenge

Challenge 2: how do you nest solutions within challenge blocks?

You can add a line with at least three colons and a solution tag.

Figures


You can use standard markdown for static figures with the following syntax:

![optional caption that appears below the figure](figure url){alt='alt text for accessibility purposes'}

Blue Carpentries hex person logo with no text.
You belong in The Carpentries!
Callout

Callout sections can highlight information.

They are sometimes used to emphasise particularly important points but are also used in some lessons to present “asides”: content that is not central to the narrative of the lesson, e.g. by providing the answer to a commonly-asked question.

Math


One of our episodes contains \(\LaTeX\) equations when describing how to create dynamic reports with {knitr}, so we now use mathjax to describe this:

$\alpha = \dfrac{1}{(1 - \beta)^2}$ becomes: \(\alpha = \dfrac{1}{(1 - \beta)^2}\)

Cool, right?

Key Points
  • Use .md files for episodes when you want static content
  • Use .Rmd files for episodes when you need to generate output
  • Run sandpaper::check_lesson() to identify any issues with your lesson
  • Run sandpaper::build_lesson() to preview your lesson locally