MSc Artificial Intelligence graduate focused on empirical AI research, multimodal evaluation, model robustness, and reproducible ML experimentation.
My current work is moving from applied ML projects toward open AI research, especially around:
- multimodal safety and harmful-content evaluation
- vision-language model evaluation
- dataset quality and leakage audits
- reproducible experiment pipelines
- failure analysis and robustness-oriented benchmarking
-
context-augmented-meme-safety-evalMultimodal safety evaluation pipeline with staged evidence construction, VLM captioning, caption selection, CLIP-family classification, zero-shot baselines, and failure-mode analysis. -
biomedical-abbreviation-nerReproducible biomedical NER benchmark with CRF, Linear-SVC, BiLSTM, BERT/RoBERTa, dataset-overlap audit, and decontaminated-train sensitivity analysis. -
vehicle-reid-experiment-tuningControlled computer-vision experiment archive on the VeRi benchmark, covering backbone comparison, augmentation, learning-rate, batch-size, and optimizer tuning.
I am especially interested in research projects where model performance needs to be interpreted carefully: cases involving noisy inputs, multimodal context, shortcut learning, dataset artefacts, or safety-relevant failure modes.