• Data Export: Download frozen candidate records for research
Benefits for Research
• Support review of candidate glycoprotein biomarkers
• Prioritize candidate relationships for external review
• Generate hypotheses for glycobiology studies
• Access versioned candidate data with explicit limitations
Disease Categories
Candidate-P-vs-U Ranking Performance
Annotation Field Coverage
Relationship Types
Protein List
Name
UniProt ID
Gene
Glyco Sites
Glycans
Status
Action
Disease-Protein Associations
Glycan Structure Library
Protein-Disease Relationships
Protein
Relationship
Disease
UniProt
Glycans
Stored provenance and rationale
Candidate Relationship Network
Layout:
Protein (circle)
Disease (diamond)
Color intensity indicates connection degree (darker = more connections). Drag to rearrange nodes.
Sequence-Based Association Ranking
Enter a protein sequence of 10–1,022 residues to obtain exploratory category-specific
literature-derived candidate-P-versus-unlabeled ranking scores from the frozen sequence-only model.
Ranking Results
Enter a protein sequence and click "Rank Sequence" to see results.
Upload FASTA File
Upload a FASTA file containing multiple protein sequences (max 100 sequences).
Validated length: 10–1,022 residues Interpretation: Not a calibrated disease probability
Download Data
Download the complete GlycoDisease database for your research.
Data reuse terms are pending author confirmation. Downloaded records are
LLM-extracted candidates and have not been independently adjudicated.
Protein-Disease Relationships
Frozen set of 39,319 candidate rows with source PMCIDs and LLM-generated rationale fields; verbatim evidence passages are unavailable.
Access candidate data programmatically. The complete machine-readable contract is available at
/openapi.json.
GET
/api/proteins
List all proteins
GET
/api/protein/{id}
Get protein details
GET
/api/relationships
List relationships (paginated)
GET
/api/diseases
List diseases with stats
POST
/api/predict
Rank sequence against candidate-P and unlabeled cohorts
GET
/api/search?q={query}
Search database
GlycoDisease Database
A literature-mined resource of candidate glycoprotein-disease relationships with
external annotations and exploratory sequence-based ranking. Candidate rows have
not been validated against an independently adjudicated gold standard.
Key Features
39,319 candidate relationship rows extracted from 4,522 PubMed Central articles
4,952 candidate protein-name records; 4,225 UniProt-enriched and 3,014 GlyGen-enriched
Literature-derived candidate-P-versus-unlabeled sequence ranking using ESM-2 embeddings