Methodology
How GLODA computes what it shows
Every record carries where it came from, when GLODA saw it, how each field was obtained, and how confident GLODA is in its own work. This page defines those terms. The live list of sources, cadences and record counts is on the coverage manifest.
Rule versions: provenance-v1 · spine-v1
Sources and cadence
GLODA reads publishers’ own channels: official APIs and feeds where they exist, official portal pages where they do not, and official bulk datasets for history. Each source is registered with its method, cadence, verification basis and, where the licence is unambiguous, its reuse terms. Sources that require a credential, publisher approval, or a browser to reach are listed on the coverage manifest as configured but not live, with the reason.
What is stored versus linked
GLODA stores the notice text a source publishes, its structured fields, the retrieval timestamp, and a content hash. Attachments and documents are linked to the source, not warehoused; a per-source terms review precedes any change to that. Raw payloads are kept immutable for a bounded retention window so a record can be re-parsed, then removed.
Provenance on every record
Publisher · source system and method · source record link and id · publication date, with whether the source stated it · last seen at the source · indexed and last updated · record hash and revision count · jurisdiction of the publisher · licence where recorded. “Last verified” in GLODA means the last time the record was seen at its source; it is not a claim of human or official verification.
Data-quality states
A state describes GLODA’s confidence in its own field, at record level and per field. It is never a judgement about the notice, the buyer or any supplier.
- Data quality: Verified
- Verified: Reserved. GLODA does not yet claim verification of any record; this state arrives with the official-domain allow-list.
- Data quality: Source-backed
- Source-backed: The field is exactly what the publisher’s own channel (API, feed, or portal page) stated, or the issuer entered it themselves.
- Data quality: Extracted
- Extracted: GLODA derived the field from the notice text with a deterministic rule (a conservative regex) or a labelled AI extraction; the original text is one click away.
- Data quality: Inferred
- Inferred: GLODA inferred the field from weak signals (keywords, an ingestion-time date, a buyer name matched from free text). Treat it as a hint.
- Data quality: Incomplete
- Incomplete: Several expected fields are missing from the source record itself.
- Data quality: Not available
- Not available: The source record does not carry this field. Missing data is never shown as a negative.
- Data quality: Needs review
- Needs review: GLODA’s own quality checks flagged the record (placeholder title, redirect page, weak summary). Check the original before relying on it.
Record-level rule, in order: review-flagged → needs review; details inferred from free text → inferred; any key field regex- or AI-extracted → extracted; three or more expected fields missing → incomplete; otherwise source-backed for official channels and issuer listings.
Extraction methods
- Official API or feed
- Official document or portal page
- Entered manually
- AI-assisted extraction — used only when a record is thin and always labelled AI; the model may not add facts beyond the notice.
- Third-party dataset — official bulk datasets (for example historical contract awards) loaded with their vintage.
- Inferred relationship — keyword sectors, regex-matched amounts, ingestion-time dates, buyer names matched from text.
Freshness
Freshness compares the record’s last update with now: fresh, recent, aging, stale. Records whose deadline has passed leave default search and are marked closed by a nightly job; records that stop appearing at a healthy source are archived. Nothing is deleted from the audit history.
Field rules for tenders
- Contract value — a structured amount from the source is source-backed; an amount matched from the text by a conservative regex is extracted (confidence 60) and marked; USD conversions use the European Central Bank rate on the publication date, or the nearest available rate, flagged as approximate.
- Sectors — published CPV / UNSPSC / NAICS codes are source-backed; GLODA sectors mapped deterministically from those codes are marked as crosswalk; keyword-only sectors are inferred and drawn de-emphasised with a tilde.
- Buyer and reference — as published, or inferred from free text and marked; an inferred buyer never merges records or scopes a procurement thread.
- Publication date — as stated by the source; when the source states none, GLODA records when it first saw the notice and says so.
- Documents and contacts — as published; document hints inferred from text are excluded from counts.
- Summary — the source’s own text; a model-generated summary replaces a missing or placeholder one and is labelled AI.
Duplicates across sources
The same opportunity published on several channels is matched conservatively (normalised title, buyer, deadline day, countries). Matches are linked, never merged: one record is elected canonical for search, every sibling stays reachable with its own provenance, and the brief lists the other sources.
Lifecycle stages
A procurement is modelled as a process with an append-only ledger of stage events, each naming its evidence (a notice, a notice revision, or an award record), the source, the method and a confidence. The current stage is the highest-ordered stage any evidenced event asserts; a submission deadline that an amendment extended re-opens a process an earlier notice had closed. Stages after award are never derived without a document that states them.
unknown → identified → planned → market_engagement → solicitation_prepared → advertised → clarification_period → amended → submission_closed → under_evaluation → awarded → contract_signed → contract_effective → in_implementation → varied → suspended → disputed → terminated → cancelled → completed → archived
Known gaps
- Several development-bank notice portals block datacenter access and are ingested through a browser lane only when a runner is available; the coverage manifest marks them dark when it is not.
- Lot-level data is stored only where a source publishes lot identifiers; GLODA never invents lots.
- Contract signature, effectiveness, variation and completion data are not yet modelled; no source in the index publishes them in a structured form.
- Financing-project links exist for World Bank project numbers; other funders’ project identifiers are wired as their sources expose them.
- Historical award datasets can be partially anonymised at the publisher (individual consultants); the record says so rather than filling the gap.
Standing statement
Data-quality states, scores, stages and flags describe GLODA’s own confidence in GLODA’s own fields. They are never findings about any organisation, buyer, supplier or individual. Only official public sanctions and exclusion lists are ever reported as such, with their source and date, and a name match is never identity.