01
Data sources
PubMed records come from NCBI E-utilities. Preprints come from the official bioRxiv API for bioRxiv and medRxiv. Interventional studies come from ClinicalTrials.gov API v2. Crossref fills missing DOI metadata and provides a narrow watched-author fallback for papers not yet indexed by PubMed. Configured RSS and Atom feeds provide headlines and short feed descriptions only; article pages are not scraped.
Each public record retains source identifiers, links, retrieval time, field attribution, and merge provenance.
02
Update schedule
A scheduled GitHub Actions workflow runs daily. Each source uses its own last-successful boundary and a seven-day overlap to catch late indexing, revisions, and corrections. A temporary failure does not advance that source or erase existing records. Bounded retries honor transient failures and fresh responses may be reused from a 24-hour local build cache; stale entries are refetched.
This is a rolling pulse rather than a permanent archive. Canonical and public data retain the latest 30 days. Publications use their publication or electronic-publication date, while clinical trials use their latest posted update.
03
Query scope and exclusions
Editable query groups cover named RNA modalities, delivery systems, and development or manufacturing. Broad discovery terms are separated from high-precision terms. The bare token “RNA” is prohibited.
PubMed discovery also monitors the configured Society for RNA Therapeutics roster. Exact ORCID is preferred. When ORCID has no works or is unavailable, person-specific PubMed and recent Crossref queries are followed by strict full-given-name and, for ambiguous identities, affiliation validation. Initials or surname overlap alone cannot label a record.
Routine transcriptomics, nontherapeutic RNA sequencing, unrelated viral epidemiology, descriptive biomarkers, and agriculture are excluded by default. Excluded candidates are retained with a reason so the decision remains auditable and rules can be changed.
04
Deduplication
DOI, PMID, NCT, source-native identifiers, and documented preprint-to-journal links are matched first. Fuzzy matching is a conservative fallback requiring a long near-identical normalized title, strong first-author identity, and compatible dates. Similar RNA terminology or a surname alone can never trigger a merge.
05
Classification
Weighted regular expressions classify modality, delivery, disease, target, stage, species, method, topic, company, and institution. Title and intervention matches carry more weight than abstract or feed-description matches. Each label stores confidence, literal matched phrases, fields, and method. Negative patterns suppress common false positives.
The Regulatory view includes formal regulatory records and strongly supported regulatory-topic classifications. A lone FDA or EMA mention in an abstract is insufficient; agency-only matches must occur in the title, while specific phrases such as “regulatory guidance” or “marketing authorization” can qualify from the abstract. Industry and Academic views use source-supplied sponsor, collaborator, and affiliation metadata. They may overlap when a development is an industry–academic collaboration.
06
Relevance scoring
Results initially appear newest first, with relevance breaking date ties. The initial view covers the 30 days ending on the dataset update date. One-click windows expose the last week, month, or all retained dates, and the From and To fields accept any custom range within the retained data. Selecting “Highest relevance” ranks the filtered result set by a deterministic 0–100 priority score.
The score is the sum of the components below. Every expanded record shows its awarded points and a plain-language reason. Classification details also retain the matched phrase, source field, confidence, and method; provenance identifies the source record. This makes the ranking inspectable and reproducible rather than a hidden model judgment.
The score is not scientific quality, clinical validity, or endorsement. Strong language in a press release receives no evidence-quality bonus.
07
Optional LLM enrichment
The default build uses no language model. When explicitly enabled, enrichment runs during the private build from public metadata only, validates structured output, and never places a credential in the website. A malformed response cannot block deterministic publication. Every build recursively scans the public site for credential-like strings before it can be published.
08
Known limitations
- Indexing and feed publication can lag the underlying event.
- Rule-based terms can miss unfamiliar products, aliases, and disease language.
- Author and organization strings are not a complete identity authority.
- Feed items represent source claims, not independently verified evidence.
- Automated classifications, merges, summaries, and scores can contain errors.
Inclusion does not represent endorsement by the maintainers or the Society for RNA Therapeutics.