A research team comprising Associate Professor FUNAKOSHI Yohei, Associate Professor YAKUSHIJIN Kimikazu, Professor MINAMI Hironobu, and Associate Professor OHJI Goh of Kobe University, together with graduate student MASUDA Genki, then-graduate student IIZUMI Shunsuke, Associate Professor OHUE Masahito, and colleagues at the Institute of Science Tokyo, has developed a new analysis method, LM-QASAS. This method combines antibody language models with B-cell receptor (BCR) repertoires(*1) collected before and after vaccination to identify and track candidate BCR sequences that are likely to have responded to an immune stimulus, without using a database of known antigen-specific antibodies. LM-QASAS extends QASAS(*2; Quantification of Antigen-Specific Antibody Sequence), a method for quantitatively evaluating immune responses to a specific antigen through antigen-receptor repertoire analysis.
When the researchers analyzed longitudinal IgG BCR repertoire data from 10 participants in SARS-CoV-2-related cohorts, the top candidates selected using an antibody language model(*3) contained a higher proportion of sequences similar to known antibodies than candidates obtained by random sampling. In addition, a pseudo-reference database(*4) constructed from nine of the participants was able to track the rise, peak, and subsequent decline of the immune response in the participant left out of database construction.
The findings suggest that LM-QASAS may make it possible to build pseudo-reference databases from the BCR repertoires of multiple individuals and use them to follow immune responses in new individuals, even for emerging infectious diseases for which little information on known antibodies has been accumulated. They also indicate that deep learning can capture biologically informative patterns in highly diverse antibody sequences. With validation across additional antigens and larger cohorts, the approach could develop into a platform for monitoring a broad range of humoral immune responses, including responses relevant to vaccines, infectious diseases, autoimmune diseases, and cancer immunity.
The research was published online in Frontiers in Immunology on July 30, 2026.

Main Points
- The original QASAS method relies on databases of known antibody sequences, which limits its application when sufficient antibody information is unavailable.
- LM-QASAS uses AbLang2 to convert IgG BCR CDR-H3 sequences into 480-dimensional representations and identifies populations whose local density rises at the response peak and then declines.
- In SARS-CoV-2-related data from 10 participants, LM-QASAS enriched sequences similar to known SARS-CoV-2 antibodies and used pseudo-reference databases built from the other nine participants to track longitudinal response patterns.
- Using clinical samples from SARS-CoV-2 vaccine recipients and convalescent patients, the study showed that the antibody language model can capture biologically informative patterns in longitudinal BCR repertoire data despite the high diversity of CDR-H3 sequences.
- Validation across more antigens and larger cohorts may enable vaccine evaluation and monitoring of humoral responses in emerging infections, autoimmune diseases, and cancer immunity.
Background of the Research
B cells carry membrane-bound forms of antibodies called B-cell receptors (BCRs) on their surface. The human body can theoretically generate as many as 100 trillion different BCRs. The complete collection of these sequences - the BCR repertoire - records an individual's immunological history, including infections, vaccinations, and autoimmune responses. Although next-generation sequencing can now read large numbers of BCR sequences at once, identifying the small subset that actually responded to a particular antigen remains difficult.
The research group's previously developed QASAS method visualizes immune responses to a specific antigen over time by comparing BCR repertoires obtained before and after vaccination with databases of known antigen-specific antibodies. This approach is useful in fields such as COVID-19 research, for which many antibody sequences have been accumulated. For emerging infectious diseases and other conditions with limited antibody information, however, a suitable reference database may not be available. The team therefore sought a method that could discover candidate responsive sequences autonomously from the internal organization and temporal dynamics of the repertoire, without relying on a database of known antigen-specific antibody sequences.
Content of the Research
1. Searching for transiently dense sequence populations in antibody 'language space'
LM-QASAS takes amino acid sequences from CDR-H3(*5), a region of the antibody heavy chain that is particularly important for antigen recognition, and processes them with the antibody language model AbLang2. Each sequence is converted into a 480-dimensional embedding(*6). Sequences with similar features learned by the model are positioned near one another in this high-dimensional language space. The researchers represented repertoires from before stimulation (Pre), the response peak (Peak), and the contraction phase (Post) as distributions in this space. They then searched for regions that were sparse at Pre, became locally dense at Peak, and decreased again at Post. Candidate sequences were identified using two complementary approaches: K-means clustering (K = 500) and k-nearest-neighbor analysis (k = 200).

Figure 1. Longitudinal changes in the BCR repertoire captured by an antibody language model (paper Figure 1C)
A representative healthy recipient of an mRNA vaccine is shown before vaccination, at the response peak, and after the peak (left to right). Background color indicates local sequence density. Red points denote sequences with at least 85% identity to CDR-H3 sequences of known SARS-CoV-2 antibodies in CoV-AbDab. At the peak, the red points accumulate in a newly formed high-density region and then decrease. The red points were used for downstream evaluation and were not used in the per-sample candidate-extraction step. Candidate extraction was performed in the original 480-dimensional embedding space, not in the two-dimensional UMAP display. Source: excerpted from Masuda et al., Frontiers in Immunology (2026), Figure 1C. © Masuda et al. (2026), CC BY 4.0.
Figure 1. Longitudinal changes in the BCR repertoire captured by an antibody language model (paper Figure 1C)
A representative healthy recipient of an mRNA vaccine is shown before vaccination, at the response peak, and after the peak (left to right). Background color indicates local sequence density. Red points denote sequences with at least 85% identity to CDR-H3 sequences of known SARS-CoV-2 antibodies in CoV-AbDab. At the peak, the red points accumulate in a newly formed high-density region and then decrease. The red points were used for downstream evaluation and were not used in the per-sample candidate-extraction step. Candidate extraction was performed in the original 480-dimensional embedding space, not in the two-dimensional UMAP display. Source: excerpted from Masuda et al., Frontiers in Immunology (2026), Figure 1C. © Masuda et al. (2026), CC BY 4.0.
In this representative case, a new high-density region appeared at the response peak. Sequences similar to known SARS-CoV-2 antibodies from the public CoV-AbDab database(*7) were found to concentrate in this region. This density and the number of red points declined after the peak, although some sequences persisted. UMAP(*8) was used only to make the pattern visible in two dimensions; density estimation and candidate extraction were performed in the information-rich 480-dimensional space. These results suggest that, despite the high diversity of CDR-H3 sequences, an antibody language model can capture biologically informative patterns beyond those detectable by exact sequence matching alone.
2. Post-extraction validation in 10 SARS-CoV-2-related participants
The validation dataset comprised longitudinal IgG BCR repertoires from four healthy recipients of an mRNA SARS-CoV-2 vaccine, three people who had recovered from COVID-19, and three people who had received an mRNA vaccine after undergoing hematopoietic stem cell transplantation. After LM-QASAS had extracted candidate sequences without using a known-antibody database, the researchers measured the proportion of those candidates with at least 85% sequence identity to known SARS-CoV-2 antibody sequences in CoV-AbDab.
For the K-means version of LM-QASAS, 26.4% of the top 50 candidates matched this criterion, compared with 1.4% of randomly sampled sequences. The corresponding values were 24.8% versus 1.1% for the top 100 candidates and 15.6% versus 1.5% for the top 300. Among the four healthy mRNA vaccine recipients, the mean match rate for the top 300 candidates was 30.8%, compared with 2.3% for random sampling. These results support the idea that the enrichment of antigen-associated sequences among candidates can be identified without reference to an external antigen-specific database.
| Extraction condition | LM-QASAS (K-means) | Random sampling |
|---|---|---|
| Top 50 candidates (all 10 participants) | 26.4% | 1.4% |
| Top 100 candidates (all 10 participants) | 24.8% | 1.1% |
| Top 300 candidates (all 10 participants) | 15.6% | 1.5% |
| Top 300 candidates (four healthy mRNA vaccine recipients) | 30.8% | 2.3% |
Table 1. Selected values from Table 1 of the paper. Values are the percentages of extracted candidate sequences with at least 85% identity to CDR-H3 sequences of known SARS-CoV-2 antibodies in CoV-AbDab. CoV-AbDab was not used in the per-sample candidate-extraction step, but was used as an external standard for downstream evaluation.
3. Tracking immune responses with pseudo-reference databases built from other participants
The researchers next excluded one of the 10 participants and used the top 300 sequences extracted from the peak repertoires of the other nine to construct a method-specific pseudo-reference database. They then analyzed the participant who had been left out and repeated this leave-one-out procedure for all 10 participants. The LM-QASAS-derived pseudo-reference databases captured longitudinal patterns that rose, peaked, and declined without using an external database of known antigen-specific antibodies. In healthy mRNA vaccine recipients, peaks were generally observed around seven days after vaccination.

Figure 2. Representative immune-response tracking using pseudo-reference databases (adapted from paper Figure 2)
Changes in sequence overlap for three healthy vaccine recipients (HV01, HV04, and HV05), normalized to 1 at the pre-vaccination time point, are shown on a logarithmic scale. Red: LM-QASAS (K-means); blue: LM-QASAS (k-NN); orange: Seq-QASAS (k-NN); green: Random; purple: CoV-AbDab. In HV04, responses associated with multiple additional vaccine doses were also tracked. Source: adapted from Figure 2 in Masuda et al., Frontiers in Immunology (2026). © Masuda et al. (2026), CC BY 4.0.
Further Developments
LM-QASAS represents a means of extracting candidate antigen-responsive sequences directly from longitudinal BCR repertoires without waiting for known antibody sequences to accumulate, and then using those candidates as a pseudo-reference database to track responses in new individuals. Potential applications include rapid monitoring of humoral immunity soon after the emergence of a new infectious disease and evaluation of vaccines against antigens for which antibody information is limited.
The study also has important limitations. Although candidate sequences showed greater overlap with known antibodies than randomly sampled sequences, random pseudo-reference databases also produced some peak-like longitudinal changes. The difference between LM-QASAS (K-means) and Seq-QASAS (k-NN) was not statistically significant at any extraction depth in this cohort. Moreover, in a separate analysis of 17 influenza vaccine recipients, clear longitudinal dynamics were not detected in 11 of them, indicating reduced sensitivity when the repertoire-level signal is weak or heterogeneous. Further methodological refinement and validation are therefore required.
The team is currently working to standardize sampling schedules that capture the pre-stimulation, peak, and contraction phases. It also aims to establish a reproducible evaluation platform using longitudinal samples from recipients of varicella-zoster virus and respiratory syncytial virus vaccines. The long-term goal is to develop a platform that can assess humoral immune responses in infectious diseases, vaccine development, autoimmune diseases, and cancer immunity without being limited by the availability of databases of known antigen-specific antibody sequences.
Explanation of Terminology
*1 B-cell receptor (BCR) repertoire: The complete collection of diverse BCR sequences present in an individual. Because particular B-cell clones increase or decrease in response to infection, vaccination, and other immune stimuli, the repertoire reflects the history of immune responses.
*2 QASAS (Quantification of Antigen-Specific Antibody Sequence): A technology developed by the research group that computationally compares a person's BCR repertoire with a database of BCR or antibody sequences known to bind a particular antigen. It quantitatively follows changes over time in the number and proportion of clones whose sequences match the database.
*3 Antibody language model: A machine-learning model trained on large numbers of antibody amino acid sequences to learn recurring patterns in those sequences. This study used AbLang2 to convert each CDR-H3 sequence into a 480-dimensional numerical representation.
*4 Pseudo-reference database: A database assembled from candidate sequences extracted by LM-QASAS or a comparison method from longitudinal repertoires. It is used in place of a known-antibody database to track immune responses in another individual.
*5 CDR-H3 (complementarity-determining region H3): The third complementarity-determining region of the antibody heavy chain. It is particularly important for antigen recognition, is extremely diverse, and provides a major marker for distinguishing BCR clones.
*6 Embedding and language space: An embedding represents the features of a sequence as a set of numbers. In this release, language space refers to a numerical space in which sequences with similar features learned by the model are positioned close together.
*7 CoV-AbDab: A public database of antibody sequences reported to bind SARS-CoV-2 and other coronaviruses. It was not used in the per-sample candidate-extraction step, but as an external standard for downstream evaluation.
*8 UMAP and KDE: UMAP is a method for reducing high-dimensional data to two dimensions for visualization. Kernel density estimation (KDE) provides a smooth visualization of point density. Both were used to display Figure 1, whereas LM-QASAS candidate extraction was performed in the 480-dimensional embedding space.
Acknowledgments
This study was supported by JSPS KAKENHI (JP25K08401 and JP26H02544), JST FOREST (JPMJFR216J), the Japan Agency for Medical Research and Development (AMED; JP25ym0126805), and the Uehara Memorial Foundation.
Original publication
Genki Masuda et al.: LM-QASAS: reference-free identification of antigen-specific sequences from the BCR repertoire using antibody language models. Frontiers in Immunology (2026). DOI: 10.3389/fimmu.2026.1844788
Release on EurekAlert!
AI tracks immune responses without relying on databases of known antibody sequences
Inquiries
For inquiries, please contact gnrl-intl-press[@]office.kobe-u.ac.jp

