About bioLLM/wiki
Last updated: 19 August 2026
bioLLM/wiki is an independent, non-commercial research project that turns recent biomedical literature into a browsable reference of the genes, proteins, chemicals, diseases, therapies and pathways researchers are publishing on right now. Every entity page summarises what the recent literature says about that entity and shows how it co-occurs with the rest of the corpus.
Where the data comes from
The underlying corpus is PubMed, queried through the US National Library of Medicine's E-utilities API. Entity identifiers are reconciled against Wikidata where a match exists. We publish bibliographic metadata and our own summaries; we do not republish publishers' full-text articles.
How an article is made
The site exists to answer one question: what is becoming important? It draws on 23,000+ biotech and medical publications, selected from the last six months of PubMed — emerging targets, therapies and pathways, ranked by recent publication activity and linked to one another through a knowledge graph. Five stages get from an abstract to an article:
- PubMed — recent abstracts are pulled through the NLM's E-utilities API into a local corpus. Only bibliographic records and abstracts, never publishers' full text.
- AI curation — a large language model reads each abstract and records the biomedical entities in it, tagged by the role each one plays in the study: background, method, target, comparator, result or conclusion. That role tag is what separates a compound a paper is actually studying from one it mentions in passing.
- Ranking — entities are ordered by how often they appear as the target of a study, not by raw mention count. The entities at the top of that ranking become the article set, which is why the list turns over as research attention moves.
- LLM wiki — an article is drafted or revised per entity from the papers that cite it, then cross-linked against every other article in the corpus so related entities reference each other.
- Knowledge graph — every entity, paper and role relationship is mirrored into a graph database. Co-occurrence between entities is computed there, and it is what the network visualisation and the "related entities" on each article are drawn from.
A sixth stage is not automated: articles are periodically re-examined and corrections are applied by hand. See the editorial policy for what that does and does not guarantee.
The pipeline runs on a schedule, so publication counts, rankings and the "recent papers" lists shift as new literature is indexed — an entity can climb or fall between visits, and that movement is the point rather than a defect. The data date in the site header is the date of the most recent run.
What this site is not
bioLLM/wiki is a literature-navigation aid for people who read primary research. It is not a medical resource, and it is not a substitute for a clinician, a pharmacist or the source papers themselves. Because the articles are drafted by language models, they can be incomplete or wrong in ways that read fluently. Please see the editorial policy and medical disclaimer before relying on anything here.
Who runs it
It is built and maintained as an independent side project by a single developer, with no institutional affiliation and no sponsorship from any pharmaceutical, biotech or publishing company. Server costs are offset by advertising. Advertisers have no input into which entities are covered or what the articles say.
Corrections and questions
If you find an error — especially a factual claim an article gets wrong — please get in touch. Correction reports are the most useful thing anyone can send us.