Authors
Folorunsho Bright Omage, Goran Neshich
Published in
GigaScience. Jul 27, 2026. Epub Jul 27, 2026.
Abstract
Membrane proteins constitute approximately 20-30% of all proteomes and represent over 60% of current drug targets. Although protein-lipid interactions play important structural and regulatory roles in membrane-associated proteins, most existing structural resources focus on identifying whether a residue lies within a membrane region, typically inferred from computational hydrophobicity-based positioning algorithms. This approach does not directly address a distinct biological question: which residues at the protein surface make direct physical contact with lipid molecules? Answering this question from experimental data is critical for understanding lipid-mediated allostery, designing lipid-mimetic therapeutics, and training accurate machine learning models for lipid binding site prediction.
We present MPLID (Membrane Protein-Lipid Interaction Database), a curated residue-level dataset comprising 4,704 membrane proteins representing 813 sequence clusters at 30% identity, 8,055,325 residues, and 80,439 experimentally validated lipid contact annotations (1.00% observed positive rate). Labels are derived exclusively from crystallized lipid molecules resolved in Protein Data Bank structures using a 4.0 Å all-atom heavy-atom distance cutoff. Because most native lipid interactions are lost during purification and crystallization, this observed rate represents a lower bound, and the non-contact class inevitably contains false negatives. The dataset uses a curated list of 117 candidate lipid identifiers across ten functional categories, including 90 PDB-derived ligand codes audited against the RCSB Chemical Component Dictionary and 27 CHARMM-style lipid identifiers encountered in cryo-EM depositions. These identifiers span phospholipids, cardiolipin, sphingolipids, sterols, fatty acids, glycerolipids, detergent mimetics (explicitly flagged), lipid A components, and CHARMM simulation nomenclature. To prevent data leakage, proteins are clustered at 30% sequence identity using MMseqs2, yielding 813 clusters partitioned into training (2,578), validation (1,051), and test (1,075) splits. Amino acid composition analysis reveals biologically consistent enrichment at lipid contact sites: tryptophan (1.88×), arginine (1.44×), glycine (1.36×), lysine (1.33×), and phenylalanine (1.23×) are enriched, while proline (0.51×), isoleucine (0.57×), and aspartate (0.59×) are depleted.
MPLID addresses a distinct biological question compared to existing resources (OPM, MemBlob, BioDolphin/PLIP): identifying residues that directly contact experimentally resolved lipid molecules rather than those positioned within computationally defined membrane boundaries. With 4,704 proteins and over 8 million annotated residues, MPLID provides the scale needed for training deep learning models for lipid contact prediction, with direct applications in structure-guided drug design and membrane protein engineering. The dataset adheres to FAIR principles and is freely available under a CC0 public domain dedication. Structurally resolved contacts represent only a subset of biological protein-lipid interactions, and MPLID is intended as an experimentally grounded resource rather than a complete catalog of lipid binding sites.
PMID:
42508035
Bibliographic data and abstract were imported from PubMed on 28 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 5
- Comments 0