Curriculum vitae · updated September 2026
PhD in Electronic Systems Engineering from Universidad Politécnica de Madrid. I study how large-scale multimodal models can account for the subjective perception of audiovisual content.
Early-Career Teaching Track Researcher (PICD) · Dept. of Electronic Engineering, ETSIT — Universidad Politécnica de Madrid · Speech Technology and Machine Learning group (GTHAU) · ivan.martinf@upm.es · ORCID 0009-0004-2769-9752 · imartinf.github.io
My research studies the subjective perception of multimedia content: what we remember from a video, how an advertisement persuades us, and what impression a person leaves when we watch them speak. I model these phenomena by adapting transformers and multimodal language models with parameter-efficient techniques, and I contrast them against biological signals —electroencephalography in particular— to bring prediction closer to mechanism.
Universidad Politécnica de Madrid, June 2026. Summa cum laude with international mention. Thesis: Modeling Subjective Perception in Multimedia with Large-Scale Multimodal Models. Supervisors: Fernando Fernández-Martínez and Manuel Gil-Martín.
Universidad Politécnica de Madrid, 2025. GPA 9.04/10 — 96.49th percentile. Distinction in the master's thesis. ETSIT 2024 Award for the best MSc thesis of the programme.
Universidad Politécnica de Madrid, 2024. GPA 9.62/10 — 93.75th percentile. Distinction in 27 ECTS. ETSIT-Empresa 2024 Award for the best academic record of the programme.
Universidad Politécnica de Madrid, 2021. GPA 8.03/10 — 93.91st percentile. Distinction in 31.5 ECTS. ETSIT-Empresa 2021 Award for the best academic record in the Audiovisual Systems track.
| Since | Position | Institution |
|---|---|---|
| 2026 | Early-Career Teaching Track Researcher (PICD) | Dept. of Electronic Engineering, ETSIT-UPM |
| 2023–2026 | Predoctoral researcher — UPM Programa Propio fellowship | GTHAU, ETSIT-UPM |
| 2024 | Research stay (Jul.–Oct.) | AI Multimedia Lab, Universitatea Politehnica din București |
| 2020–2021 | Ministry of Education Collaboration Fellowship | GTHAU, ETSIT-UPM |
147 hours taught in official BSc and MSc programmes at ETSIT-UPM between the 2023-24 and 2025-26 academic years.
| Course | Programme | Hours |
|---|---|---|
| Digital Systems II | BSc in Telecommunication Technologies and Services Engineering | 99 |
| Projects in Data and Systems Engineering | BSc in Data Systems Engineering | 21 |
| Neurodevices | MSc in Neurotechnology | 18 |
| Intelligence in Electronic Systems | MSc in Electronic Systems Engineering | 9 |
I also take part in two departmental teaching-innovation projects —DEMOSEI (2024) and DIRASEI (2023)— and co-author five teaching publications derived from them, presented at END, ICERI and GELS.
| Year | Work | Grade |
|---|---|---|
| 2026 | MSc · Deep learning models for predicting the recall of film excerpts from brain signals — Jaime León Ardyla | 9.5/10 |
| 2026 | BSc · EEG-based classification models for recall prediction in audiovisual memory tasks — Cristina Ruiz-Poveda Domínguez | 9.5/10 |
| 2025 | BSc · Multimodal language models for video memorability prediction — David Luna García | 10/10 — Distinction |
The 2026 MSc thesis led to a MediaEval 2026 Workshop contribution co-authored with the student.
Seven articles published, six of them as first or second author. Four are in the first JCR quartile, all four within the top 25% of their category. Two further manuscripts under review.
Iván Martín-Fernández, M. G. Constantin, B. Ionescu, M. Gil-Martín, F. Fernández-Martínez. ACM Transactions on Multimedia Computing, Communications and Applications, 2026. JCR 6.0 — Q1, top 25%. First and corresponding author. 10.1145/3788874
M. G. Constantin, C.-H. Demarty, C. Fosco, S. Halder, G. Healy, B. Ionescu, S. V. Luncanu, Iván Martín-Fernández et al. International Journal of Computer Vision 134(6):298, 2026. JCR 9.3 — Q1, top 25%. 10.1007/s11263-026-02880-6
M. Lobo-Alonso, Iván Martín-Fernández, I. Oropesa, R. Barra-Chicote, G. Kontaxakis, R. San-Segundo. Scientific Data, 2026. JCR 6.9 — Q1, top 25%. Second author. 10.1038/s41597-026-07745-8
Iván Martín-Fernández, S. Esteban-Romero, M. Gil-Martín, F. Fernández-Martínez. Multimedia Tools and Applications 85(1):30, 2026. JCR 3.5 — Q2. First and corresponding author. 10.1007/s11042-026-21260-3
Iván Martín-Fernández, S. Esteban-Romero, F. Fernández-Martínez, M. Gil-Martín. Sensors 25(6):1661, 2025. JCR 3.5 — Q2. First and corresponding author. 10.3390/s25061661
S. Esteban-Romero, Iván Martín-Fernández, M. Gil-Martín, F. Fernández-Martínez. Symmetry 17(8):1349, 2025. JCR 2.2 — Q2. Second author. 10.3390/sym17081349
S. Esteban-Romero, Iván Martín-Fernández, R. San-Segundo, M. Gil-Martín, F. Fernández-Martínez. Frontiers in Artificial Intelligence 9:1897208, 2026. JCR 6.7 — Q1, top 25%. Second author. 10.3389/frai.2026.1897208
Optimizing Video Transformers for Isolated Sign Language Recognition — Computer Vision and Image Understanding (Elsevier), submitted June 2026.
Evaluation of Transformer-Based EEG Image Classification Under Subject-Dependent and Subject-Independent Protocols — Applied Soft Computing (Elsevier), submitted September 2026.
Seventeen contributions to international conferences, fifteen of them as first or second author and two at a CORE A* venue. Selection:
| Year | Venue | Contribution |
|---|---|---|
| 2026 | MediaEval 2026 Workshop, Amsterdam | Three papers, one the overview of the task I co-organise |
| 2026 | MultiMedia Modeling, Prague | Attention explainability in vision-language models adapted to persuasion |
| 2026 | ICAART | Temporal landmark selection and normalisation for sign language recognition |
| 2025 | CBMI, Dublin | Size, architecture and hyperparameters in vision-language model adaptation |
| 2025 | MediaEval 2025 Workshop, Dublin | Two papers, one the overview of the task |
| 2024 | MuSe @ ACM Multimedia, Melbourne | Two papers (CORE A*) on multimodal social perception |
| 2024 | Odyssey, Quebec | Multimodal audio-language model for speech emotion recognition |
| 2024 | IberLEF @ SEPLN, Valladolid | Efficient adaptation of mono- and multimodal LLMs for speech emotion |
| 2023 | CBMI, Orléans | Memorability prediction from jointly-learnt semantic and visual features |
I am first author of the task overview papers for the 2025 and 2026 editions, which set the evaluation framework the participating teams work against.
In September 2026 I taught a two-hour workshop on generative AI for the communications departments of the eighteen clubs of the Spanish basketball league (ACB), focused on using these tools without giving up editorial judgement or each club's own identity.