Internal · in progress

Academicship

I built a bibliographic pipeline to reconcile research output and collaboration records while preserving provider evidence and protecting limited extraction quotas.

The problem

A university wants a reliable picture of its research: what its people publish and who they publish with. Bibliographic databases have that information, but not in a form you can use directly. A provider’s author profile isn’t necessarily a real, verified researcher, and a raw co-authorship link isn’t necessarily a distinct collaboration.

At UCAM I built Academicship’s data pipeline to produce that picture. The engineering goal was to keep every number traceable back to what a source actually reported, and to keep that separate from what the pipeline had reconciled.

The context

Providers disagree about IDs, affiliations and coverage. The same paper can arrive under different IDs. Citation counts depend on the source, so adding them up double-counts: two providers agreeing on a DOI doesn’t mean their citing papers can be summed.

Extraction is also expensive in a way processing isn’t. Provider quotas are limited, so a change to the parsing code shouldn’t mean downloading the same records again. OpenAlex supplies the public corpus. A Web of Science adapter is built, but live use is blocked by a shared institutional quota.

Architecture and decisions

Provider extraction

Raw yearly shards

Normalisation / identity

Quality checks

Publish snapshot

Extraction is separate from processing. Downloading from a provider writes compressed raw files, one per year. Each file is hash-checked before it’s reused, and a changed extraction plan can’t silently overwrite existing output. So I can rework parsing and reconciliation as often as I want without spending quota.

Sources stay distinguishable. The data contract keeps provider IDs, citation counts and incremental watermarks apart. Works are deduplicated by DOI, but each keeps its source IDs, and citation counts are never summed across providers. Works without a DOI keep a provider-qualified ID.

Identity is conservative. No merging authors by name alone. Observed profiles stay distinct from verified people, so a profile count never quietly becomes a headcount. A relationship only enters the published graph if both ends are valid.

Publishing is ordered and recoverable. A run publishes in a fixed sequence: snapshot, then manifest, then the CURRENT pointer, then the watermark. The watermark only moves after the snapshot is live, so extraction can’t get ahead of what’s being served. If a run is interrupted, recovery re-checks hashes and quality before moving anything forward.

Quality checks before publishing: non-null and unique IDs, numeric ranges, author and institution foreign keys, relationship endpoints, and where each citation count came from. These catch structural problems; they can’t prove every real-world identity was resolved correctly.

What it does today

An internal application runs on a retained OpenAlex snapshot. In the July 2026 snapshot that’s 6,254 works, 12,607 external co-authors, 32,016 raw relationships and 3,052 observed institutional author profiles. These are bibliographic identities in the corpus, not verified people or contacts.

The pipeline has provider adapters, shard reuse, source-aware normalisation, quality checks and recoverable publishing. There’s optional Great Expectations validation and a Spark entry point. The current runner and the Spark smoke test are capped at 100 records: that’s a bounded check of the newer processing path, not a measure of throughput on the full corpus.

Limits and what’s next

The Web of Science adapter passes its offline tests but hasn’t ingested live data, because of the shared quota. When access allows, the next step is a small live check of records and affiliations before widening extraction.

There’s no scheduled refresh yet, and I haven’t measured running costs or full-scale production runs. Before calling this a recurring pipeline I need to verify incremental loads against live provider responses and set up the refresh.

As coverage grows, identity uncertainty has to stay visible. Structural checks reject broken references; ambiguous profiles still need a person to review them.