Gianluigi Vitale presenting DriftBench at MLSys 2026 Watch the MLSys 2026 talk

ML systems researcher · Systems engineer

I study how LLM systems fail.

I build and measure the systems beneath model behavior: serving stacks, accelerator kernels, numerical paths, and the failures that appear when infrastructure changes.

Google TPU Research CloudActive since May 2026

Research built with Google TPU Research Cloud

Model bring-up, new kernels, cross-accelerator validation, and bit-level hardware characterization across TPU v4, v5e, and v6e.

Validated

DriftBench on TPU

Extended the MLSys 2026 result to TPU v5e and v6e. The predictability result held across the new numerical path, with R² from 0.72 to 0.96.

Research in progress

Bit-level TPU MXU characterization

Built a CPU simulator that reproduces real Llama-3.1-8B matrix multiplications bit-for-bit on a native pass, enabling deterministic study of TPU numerical behavior.

Collaboration

Long-context genomics with CNR Italy

A collaboration studying whether long-context DNA models encode linkage disequilibrium, combining ML systems work with genomics expertise at Italy’s National Research Council.

NeurIPS 2026 · Main trackAccepted poster · Sole author

Decomposing One Professional-Framing Pipeline: Which Components Shift LLM Safety Boundaries?

A pre-specified 2×2×2 factorial ablation of the STF professional-framing pipeline, evaluated in approximately 31,900 trials across nine LLMs. Persona adoption was the main driver of acceptance (OR = 6.5), terminology substitution the main driver of response depth (β = 0.91), and moral justification showed no detectable positive effect at the stated equivalence bounds.

02 / Academic service

Reviewing systems work from evidence to impact.

2026–2027

Program Committee Member — MLSys 2027

Reviewing submissions for the MLSys 2027 technical program.

2026

NeurIPS Ethics Reviewer

Reviewed submissions for ethical concerns, societal harms, research-integrity risks, and compliance with conference guidelines.

2026

Artifact Evaluation Committee Member — SOSP 2026

Evaluated research artifacts in memory-efficient model serving and memory-safety analysis for GPU kernels in LLM inference systems. Assessed availability, functionality, reproducibility, and consistency with the associated papers' experimental claims.

2026

Artifact Evaluation Committee Member — MLSys 2026

Evaluated four accepted submissions spanning attention-kernel optimization, ML profiling infrastructure, GPU memory disaggregation, and distributed inference. Reproduced experimental results on rented 8×H100 clusters (RunPod) and AMD MI350X infrastructure (AMD University Program), assessing availability, functionality, and consistency with each paper's claims.

03 / Research agenda

Following failures across the ML systems stack.

My work connects hardware arithmetic, kernels, serving infrastructure, and model behavior. I study where assumptions break and build evidence that makes complex AI systems easier to trust, reproduce, and deploy.

01

Systems and serving

Building and evaluating infrastructure for efficient, reliable model execution across accelerators and deployment settings.

02

Hardware and numerical behavior

Understanding how arithmetic, precision, and kernels shape correctness, performance, and reproducibility.

03

Reliability and safety

Measuring when system or context changes alter model outputs and surfacing risk before deployment.

04

ML systems across domains

Applying systems and long-context methods to domains including genomics and public-sector information access.

PyTorchvLLMSGLangTensorRT-LLMCUDACloud TPUDockerHugging Face

04 / Profile

Research grounded in deployed systems.

Experience2024—present

Systems Engineer — Archivio di Stato di Pistoia (Italian State Archive), Italian Ministry of Culture

Designed and deployed a retrieval-augmented generation system over 10,000 internal administrative documents, delivering source-cited answers to staff queries. Built an AI-assisted public consultation interface giving staff and citizens natural-language access to the full inventory database.

EducationExpected Feb. 2027

B.Sc. in Computer Engineering · Universitas Mercatorum

GPA 29.03/30 (3.79/4.0 WES-converted). Thesis: DSProfiler — static analysis and machine learning for automated detection of data-structure performance anti-patterns in Python.

Research conversations & PhD opportunities

Let’s talk about reliable ML systems.

gianluigi.vitale12@gmail.com