protein_dimension_db

Datasets with embeddings and other representations for all proteins in Uniprot/Swiss-Prot

View on GitHub

🧬🖥 Protein Dimension DB 🖥🧬

Scientific data lake with PLM embeddings, GO annotations and taxonomy representations for all proteins in Uniprot/Swiss-Prot

Current Release (2)

Proteins are sorted by length. All files contain the same sequence of proteins, to make joins and merge operations easier.

Protein Language Model Embeddings 🔢

Several models are used to create computational descriptions (embeddings) of the Swiss-Prot proteins:

Model 🤖 Vector Length 📏 Download Links (By Pooling Method) 🔗
ElnaggarLab/ankh-base 768 Mean (1.5G), Parti (3.0G), Max (1.4G), STD (1.4G)
ElnaggarLab/ankh-large 1536 Mean (2.9G), Parti (6.0G), Max (2.7G), STD (2.7G)
ElnaggarLab/ankh2-ext2 1536 Mean (2.9G), Parti (6.0G), Max (2.7G), STD (2.8G)
ElnaggarLab/ankh3-large 1536 Mean (2.9G), Parti (6.0G), Max (2.7G), STD (2.7G)
facebook/esm2_t30_150M_UR50D 640 Mean (1.3G), Parti (2.5G), Max (1.2G), STD (1.2G)
facebook/esm2_t33_650M_UR50D 1280 Mean (2.4G), Parti (5.0G), Max (2.3G), STD (2.3G)
biohub/ESMC-300M 960 Mean (1.8G), Parti (3.7G), Max (1.7G), STD (1.7G)
biohub/ESMC-600M 1152 Mean (2.2G), Parti (4.5G), Max (2.1G), STD (2.1G)
Profluent-Bio/E1-150m 768 Mean (1.5G), Parti (3.0G), Max (587M), STD (1.4G)
Profluent-Bio/E1-300m 1024 Mean (2.0G), Parti (4.0G), Max (757M), STD (1.8G)
Profluent-Bio/E1-600m 1280 Mean (2.4G), Parti (5.0G), Max (929M), STD (2.3G)
flair-bio/amplify-120m 640 Mean (1.3G), Parti (2.5G), Max (1.2G), STD (1.2G)
flair-bio/amplify-350m 960 Mean (1.9G), Parti (3.7G), Max (1.8G), STD (1.8G)

Autoencoder Embeddings and One-hot Encodings 🔢

Numerical representations of the NCBI taxon IDs and InterproScan categories of each protein. Instead of the original NCBI taxonomy tree, we use the custom taxonomy created by taxallnomy project, because it attributes the same number of parent taxa (genus, family, order…) to each species ID.

Name Description Vector Length 📏 Download Links 🔗
Interpro Autoencoder Encoding of the top 17000 most common InterproScan categories in SwissProt proteins 32 Embeddings (36M)
TaxID Autoencoder Encoding of the top 6006 most common Taxonomic IDs in SwissProt proteins 32 Embeddings (5.1M)
emb.taxa_profile_256.parquet Taxa One-Hot Encoding 256 Encodings (62M)
emb.taxa_profile_128.parquet Taxa One-Hot Encoding 128 Encodings (33M)

Protein Annotations 📚

Gene Ontology annotations of Swiss-Prot proteins, separated by evidence type. We also have made available a parsed version of the DeepLoc dataset.

Gene ontology evidence code groups included:

Gene ontology annotation columns:

DeepLoc annotation columns (subcellular locations and membrane protein types) from the Ødum et al. (2024) article:

Name Content Download Links 🔗
go.expanded.tsv.gz Simplified version of GOA. Columns: Uniprot ID, GO ID, Evidence Code, Taxon ID and Ontology 25M
go.mf.parquet Molecular Functions 14M
go.bp.parquet Biological Processes 21M
go.cc.parquet Cellular Components 9.9M
deeploc.parquet Subcellular Locations (DeepLoc dataset) 173KB
interpro.tsv InterproScan categories of SwissProt proteins 27M
taxid.tsv Taxonomic ID of each protein in SwissProt. Columns: uniprot_id, taxid, lineage 46M

Uniprot/Swiss-Prot 🔬

Name Content Download Links 🔗
ids.txt Uniprot Accession IDs 10K
sequences.swissprot.fasta Aminoacid sequences of SwissProt proteins 201M

Others

Name Content Download Links 🔗
taxallnomy.parquet Parent TaxonIDs in Taxallnomy for each NCBI taxon ID 306M
taxid.obo NCBI taxonomy graph in OBO format 2.0M
interpro.obo Interpro categories graph in OBO format 4.1M

Citation

Please cite the following work:

Bibtext:

@inproceedings{AlvesSobrinho2025ProteinDimensionDB,
  author       = {Pit{\'{a}}goras de Azevedo Alves Sobrinho and Tetsu Sakamoto and Wilfredo Blanco Figuerola},
  title        = {Protein Dimension DB: A Unified Protein Repository for Representation Learning and Functional Analysis},
  booktitle    = {BioInformatics: 21st Brazilian Congress, X-Meeting 2025, João Pessoa, Brazil, June 3–6, 2025, Proceedings},
  series       = {Lecture Notes in Computer Science},
  volume       = {16037},
  year         = {2025},
  editor       = {Marcio Dorn and Fabricio Martins Lopes},
  publisher    = {Springer Cham},
  isbn         = {978-3-032-09335-6},
  eisbn        = {978-3-032-09336-3},
  address      = {Cham, Switzerland}
}

APA reference:

Alves Sobrinho, P. de A., Sakamoto, T., & Blanco Figuerola, W. (2025). Protein Dimension DB: A unified protein repository for representation learning and functional analysis. BioInformatics: 21st Brazilian Congress, X-Meeting 2025, João Pessoa, Brazil, June 3–6, 2025, Proceedings (Lecture Notes in Computer Science, Vol. 16037). Springer Cham.

References

[1] Alex Warwick Vesztrocy and Christophe Dessimoz. “Benchmarking gene ontology function predictions using negative annotations”, Bioinformatics, 36, 2020, i210–i218, doi: 10.1093/bioinformatics/btaa466;

[2] Marius Thrane Ødum, Felix Teufel, Vineet Thumuluri, et al. “DeepLoc 2.1: multi-label membrane protein type prediction using protein language models”, Nucleic Acids Research, Volume 52, Issue W1, 5 July 2024, Pages W215–W220, doi: 10.1093/nar/gkae237;

[3] Tetsu Sakamoto and Miguel Ortega. “Taxallnomy Database”, Laboratório de Biodados, UFMG. URL;