Seminar Data-Science I (Methods)
CS 721 Master Seminar (M. Sc. Wirt. Inf., M.Sc. MMDS, Lehramt für Gymnasien)
| Lecturer | Georg Ahnert, Marlene Lutz |
| Course Format | Seminar |
| Offering | HWS/ |
| Credit Points | 4 ECTS |
| Language | English |
| Grading | Written report (40%), Report review (10%), Oral presentation (40%) and Discussion (10%) |
| Examination date | See schedule below |
| Information for Students | The course is limited to 16 participants. Please register centrally via Portal2. |
Contact
For administrative questions, please contact Georg Ahnert.

Georg Ahnert
L 15, 1–6
3. OG – Raum 322
68161 Mannheim

Marlene Lutz
L 15, 1–6
3. OG – Raum 323
68161 Mannheim
Course Information
Course Description
In this seminar, students perform scientific research, either in the form of a literature review or by conducting a small experiment, or a mixture of both, and prepare a written report about the results. Topics of interest focus around a variety of problems and tasks from the fields of Data-Science, Network Science and Text Mining.
Previous participation in the courses Network Science and Text Analytics are recommended.
Objectives
Expertise: Students will acquire a deep understanding of the research topic. They are expected to describe and summarize a topic in detail in their own words, as well as to judge the contribution of the research papers to ongoing research.
Methodological competence: Students will develop methods and skills to find relevant literature for their topic, to write a well-structured scientific paper and to present their results.
Topics
This seminar will be split into four main topic blocks. Every student will be assigned a research paper from only one of these blocks to work on. Yet, it is expected that students also actively participate in the discussion of papers from other topic blocks after they have been presented.
The topics for HWS 2026/
27 are: - Personalization X Safety in LLMs.
In this seminar, we will explore the relationship betweenpersonalization and safety in large language models (LLMs), focusing on how adapting models to individual users and their contexts can affect the safety of generated responses. We will examine challenges in defining and evaluating safety for diverse users and discuss trade-offs and approaches for developing personalized language models that remain safe while providing useful, context-sensitive interactions.- Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
- When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
- Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach
- Exploring Safety-Utility Trade-Offs in Personalized Language Models
- Data Attribution.
This seminar explores data attribution in large language models (LLMs), focusing on understanding where model knowledge, behaviors, and capabilities originate and how they can be traced back to their underlying data sources. We will discuss how identifying the sources that shape model outputs can provide insight into why models behave the way they do, while raising broader questions about influence, accountability, and the relationship between training data and emergent capabilities. - Text and Missing Data in Tabular Foundation Models.
Tabular foundation models like TabPFN and TabSTAR now rival gradient boosting on structured data, but real tables are messier than their pretraining assumes: cells go missing in non-random ways, and columns often contain free text or images that these models cannot ingest natively. This group covers how the field is closing that gap—pre-training transformers as pattern-specific imputers instead of relying on one general-purpose method, replacing the lossy PCA bottleneck for text features with learned adapters and unfrozen encoders, and building benchmarks that actually test whether multimodal signal is being used. - Modeling Human Perspectives in LLM Text-Annotations.
For subjective tasks such as toxicity or offensiveness, annotators genuinely disagree, and that variation tracks their identities and beliefs rather than being simple error—so majority-vote aggregation discards signal and bakes one perspective into the „gold“ labels. These papers trace the arc from documenting the problem to modeling it: Bayesian aggregation that separates spam from systematic disagreement, fine-tuned LLMs that turn out to learn individual annotators rather than demographic groups, and chain-of-thought traces mined for the competing rationales behind each label.- NUTMEG: Separating Signal From Noise in Annotator Disagreement
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals’ Subjective Text Perceptions
- Threading the Needle: Reweaving Chain-of-Thought Reasoning to Explain Human Label Variation
- Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
- Personalization X Safety in LLMs.
Schedule
The schedule below is preliminary, dates are subject to change.
Registration period until 07.09.2026, 23:59 via Portal2 Kick-off meeting 18.09.2026, 09:00–10:30
L 15 1-6, room 314/
315 General information Drop-out until 21.09.2026, 23:59 1st Presentation Date 16.10.2026 or 19.10.2026
L 15 1-6, room 314/
315 Presentations 2nd Presentation Date 09.11.2026 or 13.11.2026
L 15 1-6, room 314/
315 Presentations Report draft due tbd Peer review of report drafts tbd Submission deadline tbd Written Report Registration
Please register via Portal2.