Introduction to Python for Statistics in Health, Medicine and Life Sciences
Python is a powerful and flexible programming language that is increasingly used for statistical analysis in health, medicine, and life sciences. This course introduces PhD candidates and researchers to using Python for applied statistical analyses, with a strong emphasis on hands-on coding in the modern Positron environment using Notebooks and Quarto rather than programming theory.
You will learn how to work with real-world research data and widely used libraries such as pandas, seaborn, matplotlib, plotly, statsmodels, and scikit-learn. The course focuses on how to perform statistical analyses in Python (syntax, workflow, and interpretation of output), while assuming you already have basic knowledge of the underlying statistical concepts.
All software used in this course is free of charge. The only requirement is that you bring your own laptop or computer.
Fast facts
- Code: 4444
- Target group: PhD candidates and researchers in Health, Medicine and Life Sciences (with prior knowledge of basic statistical methods)
- Instruction language: English
- Minimum of 10 PhD candidates; maximum 60
- Duration: Five full days (working lectures + hands-on coding)
- Course month: September 2026
- 2 ECTS (as long as the participant follows 4 of the 5 classes and finish both assignments)
- Software: Free (Python/Positron and required packages)
- What to bring: Laptop/computer
Instructor
Dr. Benedikt Langenberg
Department of Methodology & Statistics
Phone: +31 43 388 2278
Email: benedikt.langenberg@maastrichtuniversity.nl
Goals and learning outcomes
By the end of the course, you will be able to:
- Navigate and use Positron with Quarto and Notebooks for writing and executing Python code
- Understand Python basics: syntax, data types, control structures, and functions
- Load, manipulate, and explore datasets using pandas
- Create informative visualizations with seaborn, matplotlib, and plotly
- Apply statistical methods for continuous and categorical outcomes (e.g., t-tests, linear regression, logistic regression, chi-square tests)
- Combine Python and R in one Quarto workflow using reticulate
- Build and evaluate basic machine learning models using scikit-learn
- Facilitate AI-assisted coding using DataPilot (Positron's built-in AI assistant) and Claude Code to accelerate your Python workflow
Prerequisites
- No prior programming experience is required
- Prior knowledge of statistics is expected, including methods for continuous and categorical outcomes (e.g., linear/logistic regression, t-tests, chi-square tests)
- If you want to refresh your statistical foundation, the courses “Statistics Part 1” and “Statistics Part 2” are recommended
- Familiarity with R is helpful but not required
Contents
The course combines short working lectures with guided practice in notebooks, covering:
- Getting started with Python and Quarto in Positron (syntax, variables, control flow)
- Data structures and data handling with pandas (Series, DataFrames, indexing, filtering, import/export)
- Data cleaning and transformation (missing data, reshaping, merging, summarising, preparing data for analysis)
- Data visualisation with seaborn and matplotlib (histograms, boxplots, scatterplots, bar charts, customisation)
- Statistics for continuous outcomes (t-tests, correlation, linear regression) using statsmodels
- Statistics for categorical outcomes (cross-tabulations, chi-square, logistic regression)
- A short primer on machine learning with scikit-learn (train/test split, overfitting, metrics; basic models)
- Python–R integration (using reticulate) and a final case study applying the full workflow
- AI-assisted coding with DataPilot and Claude Code (generating, explaining, and debugging code)
Teaching format
- Interactive working lectures and hands-on coding sessions
- All teaching takes place in Positron with Quarto and Notebooks, combining code, output, and explanation in one environment
Assessment
There is no formal examination. You will complete short coding exercises during the course and can optionally do a small final mini project applying what you learned to a real or simulated dataset.
Software & installation
You will use:
- Python 3.x (Anaconda recommended)
- Positron
- Packages: pandas, numpy, seaborn, matplotlib, plotly, statsmodels, scikit-learn, rpy2
- (Optional) R and the reticulate package
Installation instructions will be provided before the course starts.
Recommended resources
- Allen B. Downey — Think Stats (online)
- Wes McKinney — Python for Data Analysis
- Documentation: pandas, seaborn, scikit-learn
More information on resources will be provided before the course starts.
Course dates
| Dates | Time | Location |
|---|---|---|
| 21-09-2026 | 08:30-17:00 (the exact time is subject to change) | Location to be announced |
| 22-09-2026 | 08:30-17:00 (the exact time is subject to change) | Location to be announced |
| 23-09-2026 | 08:30-17:00 (the exact time is subject to change) | Location to be announced |
| 24-09-2026 | 08:30-17:00 (the exact time is subject to change) | Location to be announced |
| 25-09-2026 | 08:30-17:00 (the exact time is subject to change) | Location to be announced |
Information
PhD Office FHML
Available Monday, Tuesday, Thursday and Friday: 9 am – 5 pm
Phone: +31 43 38 84122
Visiting address: P. Debeyelaan 15, L. van Kleeftoren, 5.N2.030
aioonderwijs@maastrichtuniversity.nl