Methodology
A transparent, reproducible methodology
Every number in Connect51 traces back to open data and a documented method. No black boxes.
Built on OpenAlex
Connect51 is built on OpenAlex, the open catalogue of the global research record — more than 320 million scholarly works, with their authors, institutions, funders, and citations. Because the source data is open, every figure we show can be traced back to its origin.
Our analytics run on a rolling window of the most recent years, currently 2021 onward. That focus is deliberate: research strategy turns on current strength, and a five-year view gives a far sharper read on where an institution is performing now than a corpus in which decades of older output dilute the signal. Where a figure is labelled “lifetime”, it means everything currently in that window.
The citation data behind those works is not windowed. FWCI, citation counts and percentile tiers are computed by OpenAlex across its entire corpus and citation graph before we ingest them. A 2022 paper’s impact therefore reflects citations from any year and any country — including from works far outside our window. Our citation universe is OpenAlex’s entire corpus and citation graph, not a trimmed-down version of it.
What we include — and what we leave out
A figure is only as trustworthy as the works behind it, so we are explicit about what enters the corpus. We include works from across the full range of sources OpenAlex indexes — journals, repositories, conference proceedings, ebook platforms, and book series — as long as each carries the essentials for reliable analysis: a title, a detectable language, a publication date, and at least one identifiable author.
We leave out:
- Retracted works — removed entirely, so a retracted paper appears in no count, average, ranking, or benchmark.
- Paratext — front-matter such as tables of contents, indexes, cover images and mastheads, which is not research output.
- Administrative records — correction notices, retraction notices, supplementary files, peer-review reports, reference entries and grant records: the apparatus of publishing rather than research itself.
- Preprints — counted once the work is formally published, not twice. A preprint and its published article are frequently listed as two separate records, so including both would inflate output. We measure the published record.
- Incomplete records that lack a title, language, date, or any identifiable author, since they cannot be reliably placed or attributed.
We keep, and handle with care, the cases that are easy to get wrong. An uncited work counts as a genuine zero rather than being dropped — excluding it would flatter the average. A work OpenAlex has not yet been able to field-normalise still counts toward output volume; it is simply left out of impact averages, not hidden. And editorials and letters are retained rather than filtered: they are a real part of the scholarly record, and the major citation databases count them too, so removing them would make our figures harder — not easier — to compare with yours.
Field-Weighted Citation Impact (FWCI)
Raw citation counts favour large, established fields. FWCI normalises citations against the world average for the same field, year, and document type, so a value of 1.0 means world-average impact and 2.0 means twice the world average. This lets you compare a mathematician and a microbiologist on a level footing.
The h5-index (h-5)
The h5-index is the h-index measured over the last five complete years: an entity has an h5-index of h when h of its publications from that window have each been cited at least h times. It rewards recent, sustained output — consistent, well-cited work — rather than a handful of older, highly cited papers.
Standing
Standing is our composite measure of an institution’s pull in a field — combining output, citation impact, and the strength of its collaboration network. It answers a question rankings can’t: where does this institution genuinely lead?
A documented subject taxonomy
We map every work to a custom Division and Subject taxonomy so benchmarking is consistent and comparable across institutions. The full taxonomy and every calculation are documented — reproducible by anyone with access to OpenAlex.
Sustainable Development Goals (SDGs)
We tag every work against the United Nations’ 17 Sustainable Development Goals using OpenAlex’s open SDG classifier. It is a multilabel model: each goal is scored independently from 0 to 1, so a single paper can map to several goals — or to none.
We only count a goal when its score is 0.4 or higher. Below that threshold the false-positive rate climbs steeply — the model starts tagging anything that mentions “water” with clean-water goals, or “city” with sustainable-cities — which is noise rather than signal. Because SDG tags feed institutional rankings and sustainability reporting, we favour precision over recall: a missed tag is less damaging than a spurious one. The 0.4 cutoff is deliberately conservative and empirically supported — independent studies that reviewed score distributions with expert input have landed on the same value.
Patent citations
Citations from the scholarly literature tell you who is building on a researcher’s work academically. Patent citations answer a different question: whose research is being drawn on by people trying to make something. We count the patents that cite a researcher’s publications, sourced from the European Patent Office’s worldwide bibliographic data (DOCDB).
We count inventions, not paperwork. A single invention is typically filed many times over — an international application, then national filings in each country where protection is sought — and those filings share a patent family. Counting each filing separately would inflate a researcher’s number simply because their work was cited by an invention that was widely filed. We count each family once, so one invention that cites you counts as one, whether it was filed in three offices or thirty.
We publish a match only when we are confident in it. Linking a patent’s citation text to a specific paper is a matching problem, not a lookup — patent bibliographies are inconsistently formatted and often lack a DOI. Every candidate match is scored, and only high-confidence matches are published. On out-of-sample audits at that threshold, manual checking found the published matches correct in the large majority of cases; we deliberately set the bar where precision is high, accepting that some genuine citations are missed rather than reporting matches we cannot stand behind.
Every number is checkable. Each count opens into the list of patents behind it, with the specific papers each one cited. A patent-citation figure you cannot audit is one you cannot defend in a funding case or an impact statement, so we treat the drill-down as part of the metric rather than an optional extra.
What the number does not include. Office coverage is expanding, and the offices currently matched are named alongside the figure — so a zero means “no citations from these offices”, not “no patent citations anywhere”. We read the citations listed on the patent’s front page, which is where examiners and applicants record the literature an invention builds on; references appearing only in the body text are not counted. Citations from a researcher’s own patents are counted and are not separated out — building your own invention on your own research is genuine translation, but if you need the figure excluding them, it is not something we can currently isolate.
Why transparency matters
Research leaders make high-stakes decisions on this evidence. That is why our methodology is public: so you can trust the numbers, and defend them.
See the methodology in action
Book a demo and we'll walk through exactly how your numbers are calculated.
What our clients say
“Through expert consultation, rigorous analysis and access to insightful data, Connect51 developed a foundational report that provided valuable clarity on our current position and performance… The insights gained have informed our strategic thinking.”
Danielle SwanepoelExecutive Director of International Relations and Strategic Initiatives“The report combined high-quality benchmarking against national, ASEAN, and leading Asian peers with insightful analysis of research impact, collaboration, SDG performance, subject strengths, industry engagement, and publication strategies… a valuable foundation for informed decision-making.”
Kunnikar KanatharanaStrategic Advisor - University Rankings


