Alignment-free prediction of cross-reactivity in influenza A (H3N2) anticipates antigenic drift

Since its introduction in 1968, Influenza A (H3N2) has undergone continuous antigenic evolution, necessitating frequent vaccine updates. To predict antigenicity and characterize antigenic drift without multiple sequence alignments, we present FluEmbed, a computational framework that leverages protein language models. FluEmbed accurately quantified the antigenic impact of viral evolution from RNA sequences, achieving strong predictive performance against hemagglutination inhibition (HI) assay titers (Spearman correlation: ρ = 0.67-0.80). FluEmbed also outperformed sequence-distance baselines (e.g., Hamming and BLOSUM62) and phylogenetic tree-based models that require sequence alignment. Using this model, we conducted in-silico mutagenesis experiments to identify site/amino acid combinations that differentially impacted antigenicity. To systematically investigate how specific mutations influence immune escape, we defined two classes of mutations: ´constrained´, where only the most likely amino acid changes at historically mutation-prone sites were considered (thereby limiting the mutation space) and ´unconstrained´, where all possible substitutions were allowed, providing a full exploration of potential antigenic shifts. Constrained mutations often confer limited antigenic changes, whereas unconstrained mutations exhibit greater escape potential, particularly outside the dominant viral lineages. Notably, 3C.2a was the only major lineage in which constrained and unconstrained mutations showed no significant difference (p ≈ 0.9), suggesting ongoing intra-clade competition rather than inter-lineage antigenic replacement. By enabling rapid, alignment-free antigenic prediction directly from sequence data, FluEmbed could complement traditional HI assays in real-time influenza surveillance and inform vaccine strain selection decisions.