Skip to content

About

FaireduPlus: Enhancing Intersectional Fairness in Education-Focused Machine Learning Using Synthetic Data

Topics

Resources

Stars

3 stars

Watchers

2 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

Repository files navigation

FaireduSDG

Enhancing Intersectional Fairness in Educational Machine Learning through Subgroup-Aware Synthetic Data Generation

Authors:

  • Nga Pham¹
  • Minh Kha Do²
  • Quang Trung Doan³
  • Anh Nguyen-Duc⁴
  • Pham Ngoc Hung⁵

¹ Dainam University, Hanoi, Vietnam ² Latrobe University, Australia ³ FPT University, FPT Polytechnic, Hanoi, Vietnam ⁴ University of South Eastern Norway, Bø I Telemark, Norway ⁵ VNU University of Engineering and Technology, Hanoi, Vietnam

Description

FaireduSDG studies intersectional fairness (fairness across combinations of protected attributes, e.g. gender × disability) in ML models trained on educational data. It combines two stages:

  1. Synthetic data generation (generate_sdg.py) — balances underrepresented intersectional subgroups by generating synthetic rows with either an LLM (via sdgx's SingleTableGPTModel) or CTGAN.
  2. Bias mitigation and evaluation (fairedu.py) — implements the FairEDU mitigation algorithm, which decorrelates non-protected features from protected attributes via linear regression before model training, then compares fairness/performance metrics before vs. after mitigation across five classifiers (logistic regression, random forest, gradient boosting, decision tree, MLP), using aif360 metrics (Disparate Impact, Statistical Parity Difference, Average/Equal Odds Difference, etc.).

Dataset Information

Preprocessed datasets used in this project are hosted at Zenodo: https://zenodo.org/records/17933909

Test data is expected at <dataset_folder>/<dataset_name>/test_<dataset_name>.csv, and each dataset folder contains subgroup CSV splits (e.g. Gender_1_Debtor_0_Probability_1.csv) used for synthetic generation.

Supported datasets (--dataset-name):

Name Protected attributes Description
student_dropout Gender, Debtor Predicting student dropout/academic success
student_oulad gender, disability Open University Learning Analytics Dataset (OULAD)
student_performance sex, health Secondary school student performance (SP)
DNU gender, age, birthplace Dai Nam University student outcomes dataset

Third-party data sources

The CSV splits and merged files hosted on Zenodo above are preprocessed derivatives of the SP, SD, and OULAD sources cited. Refer to each original source for licensing terms before redistribution.

Code Information

  • fairedu_plus.py — CLI entry point; orchestrates generation + evaluation.
  • generate_sdg.py — synthetic data generation (LLM or CTGAN) per intersectional subgroup, and merging of generated splits.
  • fairedu.py — dataset preprocessing, the FairEDU mitigation algorithm, model training/evaluation, and results export (run_fairedu, save_results_to_file).
  • utils.py — CSV merge/split helpers and dataset-specific column definitions.

Requirements

  • Python 3.9+
  • Packages: pandas, numpy, scikit-learn, aif360, ctgan, sdgx, statsmodels
  • For LLM generation, set OPENAI_API_KEY (used by generate_sdg.generate_by_llm).

Install dependencies:

pip install pandas numpy scikit-learn aif360 ctgan sdgx statsmodels

Usage Instructions

  1. Download the dataset(s) from the Zenodo link above and place them under a local <dataset_folder>/<dataset_name>/ directory.
  2. (Optional, for LLM generation) export OPENAI_API_KEY=<your key>.
  3. Run the pipeline:
python fairedu_plus.py \
  --dataset-name student_dropout \
  --dataset-folder /path/to/original_dataset \
  --generator LLM \
  --merged-output-file-name merged_output.csv \
  --seed 42

This generates synthetic data to balance intersectional subgroups, merges it with the real training data, runs FairEDU mitigation, and writes a results file (results_<dataset_name>.csv) next to the merged training file with before/after fairness and performance metrics.

CLI options (fairedu_plus.py)

  • --dataset-name (student_dropout, student_oulad, student_performance, DNU)
  • --dataset-folder path to dataset root used to locate the test CSV
  • --generator choose LLM or CTGAN
  • --merged-output-file-name name for merged synthetic output
  • --run-splitted-file / --no-run-splitted-file choose split vs combined training files
  • --seed random seed for reproducibility

Additional examples

CTGAN without split files:

python fairedu_plus.py --dataset-name student_dropout --generator CTGAN --no-run-splitted-file

OULAD dataset with explicit dataset folder:

python fairedu_plus.py --dataset-name student_oulad --dataset-folder ./dataset

Methodology

  1. Preprocessing — dataset-specific cleaning/encoding in run_fairedu (drop nulls, encode categoricals, scale numeric features).
  2. Baseline evaluation — train each classifier on the original data; record accuracy, precision, recall, F1, and fairness metrics per protected attribute.
  3. FairEDU mitigation — for each non-protected feature, fit an OLS regression against the protected attributes; if the relationship is statistically significant (p < 0.05), subtract the fitted linear effect from the feature to decorrelate it, then drop the protected attributes.
  4. Post-mitigation evaluation — retrain each classifier on the decorrelated data and record the same metrics for comparison.
  5. Export — before/after rows per classifier are written to CSV/Excel/pickle (save_results_to_file).

Acknowledgements

This work builds upon the following open-source projects:

License & Contribution Guidelines

This project is licensed under the MIT License. See the LICENSE file for full terms.

Contributions are welcome.

About

FaireduPlus: Enhancing Intersectional Fairness in Education-Focused Machine Learning Using Synthetic Data

Topics

Resources

Stars

3 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages