P000187, Reproducible Research in Bioinformatics and Computational Biology, 4.5 Hp
Print syllabus
Syllabus
Level
Third cycle
Subject
Bioinformatics
Grading Scale
One-grade scale
The grade requirements within the course grading system are set out in specific criteria. These criteria must be available by the course start at the latest.
Course language
English
Entry Requirements
Admitted to a postgraduate program in animal science, biology, veterinary medicine, food science, nutrition, nursing, or related subjects, or to a residency program in veterinary science. Alternatively, admitted to a postgraduate program in Bioinformatics or Biological field with a significant computational aspect.
Objectives
On completion of the course, the student will be able to:
- Understand the FAIR principles. Explain what reproducibility means in scientific research and why it is essential. Apply the principles to biological datasets. Evaluate their own analysis workflow in these terms.
- Identify suitable public databases for accessing and storing given types of biological data.
- Use tools in R to enhance reproducible research. Apply basic principles of reproducible analysis in R, including organizing code and results.
- Understand Unix philosophy. Be able to use navigate its file system, install and update software, log on to and use remote servers.
- Use VSCode or other development environments. Be able to write simple programs for scientific computing.
- Set up and maintain environments using tools like Pixi and Apptainer
- Understand and use Git to collaboratively develop, document, maintain, and distribute code.
- Understand and use common pipelining tools such as Nextflow and Nf-core. Design and run workflow pipelines.
- Apply the tools and principles to their own data
Content
In Module 1:
Students will learn the philosophy of R, as well as commonly used functions, how to structure code, and how to import libraries. R markdown, making R packages, functions.
Students will learn the FAIR principles in detail.
Students will learn the principle and practice of version control with Git, including git in R studio. They will learn to parallelize R projects.
Students will learn to use the Linux command line. They will be able to navigate the file system, log on to remote servers, and use basic commands. They will also learn to use NAISS High Performance Computing resources.
Students will be introduced to pipelining as a means of packaging their data and methods.
Module 2 will cover:
Version control tools. Students will learn the philosophy of version control with Git as an example. They will demonstrate understanding of basic concepts with a quiz. They will install Git on their own laptops. We will hold an interactive session in which they will collaborate on a single document which will be turned into a coherent whole.
They will create a Quarto blog, publish with github actions
Students will use Nextflow and nf-core to pipeline common bioinformatics tools as is done in bioinformatics publications to ensure reproducibility.
Students will learn about containers such as Docker and Singularity. They will learn hands-on how to use them to package their code and data for reuse under the FAIR principles.
Students will learn visualization using R and ggplot.
Final assignment is a Quarto blog explaining their Reproducible Research learning journey, version controlled using Git.
Students will review the FAIR principles in brief, and discuss how to organize computational projects.
Examination Formats and Requirements for Passing the Course
Examination will be by quizzes and written assignments.
Evaluation will be based on theoretical understanding, technical<br>
performance, and ability to interpret results.
Responsible Department/Equivalent
Department of Animal Biosciences
Supplementary information
Other Information
The course is proposed to run jointly with MedBioInfo. The course merges Reproducible Research course (geared towards VH PhD students) as Module 1, with the MedBioInfo Applied Bioinformatics course (geared towards Bioinformatics PhD students) as Module 2.
Module 1 (to be offered only through GS-VMAS) improves the students’ computational fluency, to the point that they will be able to complete the coursework of Module 2 alongside MedBioInfo students.
As students from all over Sweden will be in attendance, we intend to run a short excursion, e.g. to Lövsta or SciLifeLab. We also typically hold a course dinner for students to socialize.