Skip to main content

PhD project by Ricardo Muñoz Sánchez

Fairness in the Age of Automation: Auditing Second Language Assessment Systems for Bias and Fairness

Supervisors: Elena Volodina and Simon Dobnik

Here you can find a bit more information about my PhD topic and things related to it, including information about previous seminars and the papers that I plan to include in it. For more information about my project, feel free to check out my website.

Thesis Abstract

The past few decades have seen algorithmic automation permeate multiple aspects of life. Language asessment is no exception to this trend: several standardised testing organisations have already integrated such systems into their workflows. However, AI models are known for their lack of transparency and for reproducing biases in the data. This clashes with the concept of fairness in language assessment, which enshrines ideals such as explainability and nondiscrimination. We must therefore ask ourselves whether it is possible to have fair systems for automated language assessment or whether that concept is paradoxical by nature.

In this thesis, we explore this question from multiple facets: whether these systems tend to commit mistakes when classifying texts, which linguisitc aspects the underlying models rely on, and whether they make spurious correlations based on aspects that should not be relevant when determining language proficiency. We focus on systems that assign proficiency levels what align with those from the Common European Frame of Reference (CEFR). While we mostly center around Swedish learner essays, we also consider texts written by learners of English and French.

The backbone of this thesis comprises five articles. The first two explore which features of learner language are encoded by language models for CEFR level classification. The third and fourth articles focus on whether these systems present ethnicity or gender-related biases based on names appearing within the texts of learner essays. The final article rounds out this thesis by showing how the knowledge and insights gathered from the previous studies can be ported to another domain, namely reporting and social biases.

We conclude that even though CEFR level classification systems appear to rely on learner language characteristics rather than non-relevant ones, the underlying models show much lower performance than what would be acceptable for any high stakes application. Furthermore, there are multiple societal and ethical considerations regarding automation of language assessment in real-world settings. Due to these reasons, the development of systems for automated language assessment should be approached in a critical manner.

Seminars and Presentations

  • Final Seminar
    • Date: March 25, 2026
    • Time: 13.15-15:00
    • Room: C350
    • Opponent: Beáta Megyesi
  • Halfway Seminar
    • Title: "From Algorithms to Classrooms: NLP for Second Language Learning as a Case Study for Bias and Fairness in AI"
    • Date: November 18th, 2024
    • You can find the slides for the presentation here
  • Idea Seminar
    • Title: "Using the Flow of Information to Detect False News"
    • Date: January 23rd, 2023
    • You can find the slides for the presentation here

Papers Included in the Thesis

  • Tom Södahl Bladsjö, Ricardo Muñoz Sánchez. “Introducing MARB — A Dataset for Studying the Social Dimensions of Reporting Bias in Language Models”. 6th Workshop on Gender Bias in Natural Language Processing, co-located with ACL 2025. (link)

  • Ricardo Muñoz Sánchez, Simon Dobnik, Elena Volodina. “Harnessing GPT to Study Second Language Learner Essays: Can We Use Perplexity to Determine Linguistic Competence?”. BEA 2024 Workshop, co-located with NAACL 2024. (link)

  • Ricardo Muñoz Sánchez, David Alfter, Simon Dobnik, Maria Irena Szawerna, Elena Volodina. “Jingle BERT, Jingle BERT, Frozen All the Way: Freezing Layers to Identify CEFR Levels of Second Language Learners Using BERT”. NLP4CALL 2024. (link)

  • Ricardo Muñoz Sánchez, Simon Dobnik, Maria Irena Szawerna, Therese Lindström Tiedemann, Elena Volodina. “Did the Names I Used within My Essay Affect My Score? Diagnosing Name Biases in Automated Essay Scoring”. CALD-Pseudo Workshop, co-located with EACL 2024. (link)

  • Ricardo Muñoz Sánchez, Kostiantyn Kucher, Filip Kayar, Elena Volodina. "ABC Diagnostics: A Web-Based Tool for Explainable CEFR Classification". In review.