Friday, October 2, 2026

Column · @relatedcommand577

Inside MCP for Wikidata Candidate Search Limits

Filed by @relatedcommand577

Candidate search limits sound like a small implementation detail until you have to trust the output. Then they become one of the most important design choices in the whole system.

That is especially true in entity resolution work, where the difference between "probably right" and "provably inspectable" decides whether a match can be used downstream at all. The open source project often described as Wikidata + Google Knowledge Graph MCP takes a very specific stance here. It does not try to flood an agent with every possible match. Instead, it uses bounded search, returning three candidates by default and up to five at most. That decision shapes the behavior of the server, the quality of its evidence, and the way a human or an agent can reason about uncertainty.

If you spend enough time around record linkage, catalog cleanup, metadata harmonization, or knowledge graph enrichment, this design choice feels less restrictive than disciplined. Large candidate sets look powerful on paper. In practice, they often create noise, false confidence, and brittle automation. A narrow candidate window forces better ranking, clearer evidence, and more honest outcomes.

Why bounded candidate search matters more than it first appears

Most search systems are built around abundance. If you type a person, company, or place name, the system returns pages of results. That makes sense for interactive search. It works less well for deterministic resolution, especially inside an MCP workflow where an agent is expected to make a grounded decision.

The stated purpose of this MCP server is not generic discovery. It is to let AI agents search Wikidata, read selected facts, and link local records to Wikidata QIDs with inspectable evidence and explicit uncertainty when evidence is insufficient. That last clause matters. If uncertainty is part of the product, then the search phase cannot be a loose bag of vaguely relevant entities. It has to be constrained enough that every candidate can be examined, compared, and, if needed, rejected.

Returning three candidates by default creates a working set that can actually be reasoned over. A human reviewer can inspect it. An agent can compare fields and supporting facts without getting lost in a sea of edge matches. Raising the maximum only to five keeps the same discipline intact. You still have a shortlist, not a search dump.

I have seen the opposite approach in internal data projects, where teams ask for "all possible matches" because they do not want to miss anything. They usually get what they asked for: twenty near duplicates, several namesakes, some weak lexical matches, and one correct entity hidden in the middle. From there, either a human spends too much time on review, or automation starts leaning on shallow signals because there is https://context7.com/gitlab_revanalex/wikidata-google-knowledge-mcp simply too much candidate mass to evaluate carefully. Bounded search is a way of refusing that failure mode up front.

What this MCP server is actually trying to do

The project sits in a practical middle ground between search and resolution. It is an open source MCP server and CLI, published under MIT, and it is designed to work in MCP clients such as Claude Code, Cursor, and Codex. It can query Wikidata without requiring an account or API key, while the Google Knowledge Graph Search API is optional rather than mandatory.

That alone says something about its intended use. This is not a heavyweight reconciliation platform with endless knobs. It is read only, it does not edit Wikidata, Google, or user data, and it is explicit about not being official Wikimedia or Google software. Those boundaries are healthy. They keep the server focused on search, evidence retrieval, and resolution outcomes rather than drifting into a half managed curation tool.

For broader context, Wikidata itself has documented MCP support as a way for language models to explore and query Wikidata programmatically through standard interfaces. This project narrows the aperture further. It is not just about querying the graph. It is about querying it in a way that supports candidate ranking and defensible linking.

The phrase MCP for Wikidata fits that use case well, but the more precise framing is that this is MCP for google knowledge graph and wikidata in the service of entity resolution. The Google side is optional and carefully scoped. The Wikidata side is central.

The quiet logic behind three candidates, not thirty

A limit of three by default can look conservative if you come from search engineering. It looks normal if you come from adjudication.

Three candidates forces the system to answer a hard question: which options are actually worth serious consideration? That is not merely a user experience choice. It is a signal that the ranker must separate plausible identities from background clutter before anything reaches the caller.

There is also a practical cognitive reason for this threshold. Once you move beyond a small handful of options, comparison quality drops. Reviewers begin to skim. Agents start summarizing rather than examining. Fine distinctions, such as a mismatched occupation, a wrong country, or a slightly off temporal fact, get blurred. A shortlist preserves attention.

The cap at five is just as telling. It leaves room for difficult name collisions without opening the floodgates. Think of common personal names, organizations with similar branding, or places that share names across countries. In those cases, a hard cap of one would be reckless, but a cap of fifty would be irresponsible. Five acknowledges ambiguity while preserving Wikidata MCP bounded reasoning.

This is the core trade off:

| Choice | What you gain | What you risk | |---|---|---| | Large candidate sets | Broad recall and more raw options | Review fatigue, noisy ranking, weak evidence handling | | Bounded candidate sets | Focused comparison and clearer outcomes | Some long tail candidates may never surface |

That trade off is not abstract. It determines whether a resolver can honestly say "I do not know" instead of stretching to fit a bad match into the workflow.

Resolution quality depends on what happens after search

Candidate limits matter because they feed a deterministic resolution layer. The project documents explicit outcomes such as AUTO_MATCH, HOLD, AMBIGUOUS, and NO_CANDIDATE. Those labels are more than status codes. They describe the contract between search and decision.

A system that returns a sprawling result set often ends up blurring these outcomes. When there are twelve weakly plausible entities, "ambiguous" starts to mean "we did not rank carefully enough." When there are no strong options, some systems still promote the least bad one because the workflow expects a candidate. That is how bad identifiers get embedded into local databases.

Here, bounded search and deterministic outcomes reinforce each other. If the shortlist is tight and the evidence is inspectable, AUTO_MATCH can remain meaningful. If evidence is insufficient, HOLD or AMBIGUOUS can be used without embarrassment. If nothing clears the bar, NO_CANDIDATE is a valid end state rather than a failure of the software.

That design also reflects lived reality in reconciliation work. Not every record should resolve. Some source data is too sparse. Some names are too common. Some entities have changed names, split, merged, or are simply underdescribed in the target graph. A system that cannot stop itself from matching is more dangerous than one that occasionally declines to decide.

Evidence is where the limits pay off

The project supports selected fact retrieval, including ranks, qualifiers, and references on request. This is where the small candidate window becomes genuinely useful.

Imagine trying to compare twenty entities using full claim context. You would spend most of your time triaging what not to read. With three candidates, the opposite happens. You can afford to look at richer evidence for each one. Rank information matters because not all statements in Wikidata carry the same standing. Qualifiers matter because a statement without context can mislead. References matter because provenance changes how much weight a fact deserves during a tie break.

This is one of the strongest arguments for bounded search in knowledge graph resolution. If you want factual inspection, you need a manageable number of entities. Otherwise, the temptation is to collapse everything back into string similarity and popularity cues.

In my own experience, many mistaken links happen not because the right entity was absent, but because no one looked one layer deeper. A birthplace was wrong. A corporate parent did not line up. An occupation had a different timeframe than the local record implied. Rich evidence only helps when the candidate set is small enough to inspect.

The MCP tools reveal the intended workflow

The documented toolset includes kg_search, kg_entity, kg_related, kg_resolve, and kg_status. There is also CLI support for batch work and evidence export. Even without speculating about internals, those tool names tell a coherent story.

You search for candidates. You inspect an entity. You look at related context where needed. You run a resolver that emits an explicit outcome. You check status. In the CLI, you can scale that process across many records and export the supporting evidence.

That workflow would degrade badly if candidate retrieval were open ended. kg_search would become a firehose, and every downstream tool would inherit the sprawl. By keeping the front end bounded, the rest of the chain remains legible.

There is an underrated operational advantage here as well. Bounded outputs are easier to audit. If a batch process resolves thousands of local records to QIDs, investigators do not want to reconstruct decisions from giant result sets. They want to see the shortlist, the chosen entity, and the evidence that tipped the balance. The smaller the search envelope, the easier it is to reconstruct why the resolver behaved as it did.

Where Google Knowledge Graph fits, and where it does not

The project allows an optional Google cross check using exact identifier joins. Specifically, it documents /m/ joins for Wikidata property P646 and /g/ joins for P2671. That is a sensible use of Google data because it avoids vague semantic blending. The match point is an exact external identifier relationship, not a hand wavy "these descriptions look similar."

Just as important is the caveat: agreement between Google and Wikidata is treated as provider concordance, not proof of identity. That is the right posture. Two providers lining up can increase confidence that you are looking at the same conceptual entity, but it does not independently prove correctness in every case.

This is where the phrase MCP for google knowledge graph deserves careful handling. Some readers hear it and assume broad graph fusion, enrichment, or hidden authority signals. That would be the wrong expectation. The project is explicit that it is not an export of the Google Knowledge Graph. The Google API is optional. The server remains grounded in bounded search and inspectable evidence, with provider agreement used as one signal among others.

A lot of reconciliation tools get into trouble by overvaluing concordance. If two sources agree on an identifier, reviewers relax too quickly. But concordance can sometimes reflect a shared historical mistake or a broad category alignment rather than the exact local record you need. This server's framing avoids overstating what cross provider agreement means.

Limits are not only about performance, they are about behavior

People often assume search limits exist to save tokens, bandwidth, or latency. Those benefits are real, but they are not the interesting part here. The deeper function of a candidate limit is behavioral. It shapes what kind of system you are building.

With a bounded result set, the resolver is pushed toward clarity. It must rank well. It must explain itself. It must preserve uncertainty when the evidence is thin. It cannot hide behind abundance.

With an unbounded result set, a different behavior emerges. Ranking quality can slip because "the correct answer is in there somewhere." Evidence can become shallow because there are too many entities to inspect. Uncertainty gets diffused into volume rather than made explicit. The system feels comprehensive while actually becoming harder to trust.

That distinction matters a great deal in enterprise settings. Teams often measure success by throughput first and precision second, until a few bad links pollute reporting, analytics, or customer records. At that point, they rediscover that a narrower and more honest candidate pool would have saved time rather than cost time.

Edge cases where the limit will feel tight

A bounded candidate search is not magic. There are situations where three, or even five, will feel tight.

One obvious case is heavy name collision. A common person name can map to multiple notable individuals with similar occupations or time periods. Another is sparse local metadata. If the input record contains only a name and no geography, dates, or type hints, even a strong ranker has less to work with. A third is cases where the correct Wikidata item is weakly connected to the most obvious lexical form, perhaps because of transliteration differences or naming conventions.

In those moments, users who are used to broad search may feel the system is withholding possibilities. Sometimes it is. That is the cost side of bounded retrieval. But there is a reason the project pairs small candidate sets with explicit outcomes like HOLD, AMBIGUOUS, and NO_CANDIDATE. When the shortlist does not support a clean decision, the system is expected to say so.

That can be frustrating in the short term, especially if you want every row in a spreadsheet resolved by the end of the day. In the long term, it is usually the safer discipline.

How to judge whether the limit is helping or hurting

When teams evaluate a resolver like this, they often ask the wrong first question. They ask whether the system found the answer for a handful of easy examples. A better question is whether the system behaves well on hard examples.

Here are the signals I would watch when assessing bounded candidate search in practice:

  1. How often does the shortlist contain only genuinely plausible entities, rather than one plausible entity padded with lexical debris?
  2. When the resolver returns AUTO_MATCH, can a reviewer inspect the evidence and understand why the match was justified?
  3. When evidence is thin, does the system hold back with HOLD or AMBIGUOUS, or does it force a decision?
  4. In batch work, can the exported evidence support later auditing without reconstructing the whole search process?
  5. When Google cross checks are used, are they treated as corroboration rather than final proof?

Those questions reveal much more than raw match volume. They test whether the candidate limit is acting as a quality control mechanism or as an accidental blindfold.

What this means for MCP workflows in the real world

The phrase MCP for wikidata might sound niche, but the workflow is broadly relevant. More teams are putting language models in front of operational knowledge tasks, and entity linking is one of the first places where optimism meets friction. The model can reason fluently, but if the retrieval layer is sloppy, the reasoning rests on weak ground.

That is why bounded search belongs in the same conversation as prompt quality, schema design, and audit trails. A model that receives three well ranked candidates with selected facts, ranks, qualifiers, and references on request has a fair chance to produce a careful resolution path. A model that receives twenty noisy candidates will often sound confident while doing shallow comparison.

The same logic applies to MCP for google knowledge graph when the Google API is present. The value is not that another provider exists. The value is that the provider can be brought in through exact identifier joins and interpreted with restraint. Done that way, optional cross checks strengthen the evidence model without turning the workflow into a black box.

There is also a healthy product boundary here. Because the server is read only and does not edit Wikidata or Google, it can focus on retrieval integrity. That separation matters. Search, evidence gathering, and resolution are hard enough without mixing in write back logic or curation side effects.

A disciplined limit is often a sign of mature design

There is a tendency in tooling to equate bigger with better. More connectors, more results, more flexibility, more automation. Sometimes that is right. In entity resolution, it often is not.

This project's bounded search approach signals a mature understanding of the task. The server is not pretending that search volume equals certainty. It is not disguising ambiguity behind long result lists. It is not overstating what optional Google concordance can prove. It is trying to keep the candidate pool small enough that evidence remains inspectable and outcomes remain explicit.

That is why the candidate search limit is not a minor footnote. It is the architectural expression of the project's trust model.

If you are evaluating tools in this space, pay close attention to choices like this. The most useful resolver is rarely the one that returns the most. It is the one that knows when to stop, what to show, and how to admit doubt.

— 30 —