Authors
li, y., Gao, B., Liu, W., Xie, J., Tsui, T. Y., Zhang, Y., Wang, Y., Fu, T., Li, Y.
Abstract
Recent work applies large language models (LLMs) to single-cell data, but the cell usually reaches the model as a short ranked list of genes. This list drops much of what defines a cell, so the model reasons from a partial view. Yet the information is not lost in the measurement, only in the text. Here, we introduce Cell-Lens, a training-free structured representation of the measured cell. It writes the cell as typed blocks, from paired measurements to marker-supported programs. On CellVerse, SOAR, CellPuzzles, and SC-Arena, across five model families, Cell-Lens improves performance by up to 27.1 points. It also narrows the model gap and reduces API-reported reasoning-token usage. The largest gains occur when a defining RNA marker falls out of the ranked list. In these cases, the paired protein data still capture the corresponding signal. Shuffling the added blocks removes the gain. The gain persists after selected label-associated proteins are removed. These controls link the improvement to cell-matched biological content. We also build and release CellSpectrum, a broader benchmark spanning eight sources and seven tasks with paired modalities. The gains hold there across tasks, tissues, and modalities. These results show that a better cell representation can improve single-cell reasoning without retraining the model. Code and CellSpectrum are publicly available at https://github.com/LiZaiyuan0619/Cell-Lens.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 25 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 2
- Comments 0