Solutions
Retention time drift across batches
The same compound elutes at slightly different times from run to run. Over a large study, that drift breaks the assumption that a feature at a given time is the same molecule everywhere, and alignment quietly fails.
The problem
Retention time drift across batches
Retention time drift is the gradual shift in when a compound elutes across injections. Column aging, mobile phase preparation, temperature, and pressure all move peaks by seconds. Within a single plate the shift is small. Across hundreds of samples run over weeks it accumulates, so a feature that eluted at one time early in the study elutes noticeably later by the end. Any step that matches features by a fixed time window will then match the wrong things or fail to match at all.
Why it happens
Where conventional workflows produce it
Most alignment methods pick anchor peaks and warp each sample toward a reference. That works when drift is modest and anchors are abundant and unambiguous. In large cohorts the anchors themselves drift, low abundance features have no stable reference, and injecting in batches introduces steps rather than a smooth trend. Warping to a single reference sample also bakes that sample's idiosyncrasies into the whole study, and it degrades when the cohort is heterogeneous.
How Metablify addresses it
Evidence from the whole cohort
Metablify aligns by building a cohort wide model of how each feature moves, using the many injections as mutual references rather than trusting one reference sample. Consistent mass and elution behavior across the cohort identify which peaks are the same analyte even when absolute time has shifted. Because the model spans the whole study, batch steps and slow trends are handled together, and low abundance features borrow stability from the cohort instead of needing their own anchor.
What changes
What you see in the output
Features match correctly from the first plate to the last, missing values that were really alignment failures reappear, and batch structure stops dominating the first components of a multivariate model. The study reads as one experiment rather than a set of loosely joined batches.
Honest limits
What this does not do
Alignment cannot rescue a run where the gradient failed or where a compound moved outside the acquired window. If two compounds swap elution order under drift and share mass, separating them requires an orthogonal dimension. Metablify aligns what was measured; it does not reconstruct chromatography that was lost.
What matters
Where this makes a difference
Drift accumulates across a study
Shifts that are trivial within a plate add up across weeks of acquisition, so a fixed time window that worked early in the study quietly matches the wrong features later.
The cohort as its own reference
Rather than warping every sample toward one reference run, Metablify models movement across all injections, so low abundance features borrow stability from the whole cohort.
Missing values reappear
Many apparent missing values are alignment failures, not true absence. Correct alignment recovers them and changes the statistics that follow.
Questions
Common questions
How is this different from normalization?
Normalization adjusts intensities to make samples comparable. Alignment decides which peaks across samples are the same feature. Drift is an alignment problem. Normalizing intensities does nothing to fix features that were matched to the wrong molecule.
Do I need QC injections for this to work?
Quality control injections help characterize drift and are good practice, but the cohort itself provides most of the evidence. Metablify uses agreement across all injections rather than depending on a small set of reference runs.
Sources
- Tautenhahn et al., retention time correction, BMC Bioinformatics (2008) ↗Nonlinear retention time alignment approach.
- Dunn et al., large scale metabolomics practice, Nature Protocols (2011) ↗Batch structure and drift in large studies.
Keep reading
Related
Prove it on your own data
Send a limited set of your existing LC/MS data and see how many additional real mass features are recovered against your current output.
Send a multi batch dataset and see how many features align across the full study rather than within single batches.