Skip to article
Skip to main content

Land small, prove fast, expand - Explore the Squirro AI Agent Catalog – Download Now

Blog

R&D Knowledge Management: How to Avoid Costly Duplicate Research

Researcher at a crowded lab bench lowering a multichannel pipette into a half-filled 96 well plate, racks of tubes beside it

Key Takeaways

  • Your project registry catches live overlaps but drops closed work, which is where repeated experiments come from.
  • A failure report filed under an experiment ID is unreachable by a search for the failure itself, because the root cause sits in writing rather than in a dedicated field.
  • Checking a research plan against prior work needs the archive to be searchable by intent, not by identifier.
  • A knowledge graph sharpens expertise location and cross-site discovery, but deployment does depend on one being in place.

Every R&D organization writes up its failures. The report goes into the document system, and in most organizations that is the last time anyone opens it.

The archive knows that record by what it was: an experiment ID, a material system, an author, a site, a close date. The researcher who would benefit most from it two years later knows something different. They know what they are about to try. They have a plan with materials in it, a process, and an objective written in their own words. Because the context is so different, nothing in that plan matches how the earlier work was filed

That mismatch is a predominant cause behind duplicated research. Project registries, quarterly cross-site reviews, and innovation calls all assume that somebody in the room will notice the overlap. When no one knows that the new plan exists, they fail.

Our R&D Knowledge Discovery and Deduplication Agent, featured in our growing AI Agent Catalog, scores a research plan against an indexed archive of prior work and connects a question to the people and projects behind the documents that answer it, all before its budget is committed. What makes that possible is how documents are processed as they enter the index.

Why Coordination Fixes Don't Close the Gap

Coordination mechanisms work on the projects people already know about. As we just saw, duplication happens when nobody remembers the past work.

Coordination mechanism

What it catches

What it leaves open

Project registry or portfolio review

Two live projects with overlapping scope

Closed work. Completed and cancelled projects drop off the registry, and failures are the first to go

Cross-site technical forum

Overlaps between teams whose members attend

Sites, functions, and contractors outside the invite list

Stage-gate review

Weak justification in a proposal

Prior work the proposer never found, which the reviewers also do not know about

Skills database or expert directory

Who to ask, if you already know what to ask

The question you do not know to ask, phrased in vocabulary the directory does not hold

Departure handover

What the leaver remembers to write down

Everything they knew implicitly, and every conclusion filed under a project name nobody recognizes


Departures, the leading cause for rampant corporate amnesia, make things even more challenging. An expert leaves, the handover document is thorough by the standards of handover documents, and two years later nobody can reconstruct why a promising route was abandoned. The conclusion exists, but sits somewhere in a report filed under an experiment ID.

What a Failure Report Would Need to Contain to Be Findable

For a failed experiment to surface when it matters, the failure mode, the cost, the duration, and the recommendation need to be fields included on the record. So, let’s go ahead and take a step back, walking through the process end to end.

A typical failed-experiment report arrives as a PDF. The title carries the experiment ID and the material system. The metadata carries an author, a site, and a close date. The root is described in writing in a discussion section, along with the recommendation that came out of it.

An index built from that document's text and filename can tell you the experiment happened. Where it fails is that it doesn’t tell you that, in the experiments, the epoxy matrix degraded above a particular filler loading, which is the only part the next researcher needs.

Extraction at ingest changes the record. Outcome, failure mode, cost, duration, materials, and process steps are pulled from the document and indexed alongside its full text, so the same PDF becomes searchable by the thing that went wrong. Document type, site, outcome, and year become filters. A researcher can narrow a search to failed experiments from one site since 2022. An administrator can use the same fields the other way, to define which slice of the index a given team is allowed to query at all.

Scoring a Research Plan before the Budget Is Committed

Next, the researcher describes the experiment that they are about to run, in their own words, with the materials and the process steps. The system scores that description against the indexed metadata of everything in the archive and returns the closest prior work above a similarity threshold, with the outcome of each match attached.

Because the scoring runs on indexed fields, the same plan returns the same number every time it is checked. In other words, it’s deterministic, which is essential for a result that can sway a budget decision.

This determinism is also where chemical R&D search departs from general enterprise search applied to chemical R&D. Enterprise search matches a query to documents. A plan check compares a structured description of intended work against structured descriptions of completed work, and ranks by how much of the intent overlaps.

We’ve built an R&D Research Deduplication Agent demo to showcase this agent. Its archive comprises 42 documents across four sites. There’s a plan for a solvent-free epoxy barrier coating on PET film, plasma pre-treated, with an alumina nanofiller, scores 88% against a closed experiment from another site. That experiment ran 28 weeks, carried a recorded cost, and failed on thermal degradation above a certain filler loading. Because root cause and recommendation are indexed fields on the record, the match arrives with the post-mortem attached rather than simply pointing at a PDF.

The same check surfaces two other things: a successful variant on a different film from the same lab, and a second-phase retry running right now at a third site. Avoiding the repeat is the obvious value, as is joining work already in motion.

When the Prior Work Doesn't Share Your Vocabulary

Oftentimes, the most relevant prior work is written in language that has little in common with the question. A knowledge graph is what connects the two, so let’s assume that one is in place for this section.

Is a semantic layer essential? No, but it helps: Grounded retrieval over an indexed archive is the baseline, and it carries the duplication check on its own. Squirro Graphite is the layer above it, building and managing the taxonomies, thesauri, glossaries, and enterprise knowledge graphs that link people to projects to documents, on Linked Data and Semantic Web standards, importing existing vocabularies in SKOS, OWL, RDF, JSON-LD, and CSV. Organizations with an established vocabulary get more out of the system sooner.

What a knowledge graph adds is reachability. A question about barrier coatings returns a retry at another site whose report never uses those specific words in the question, because the graph connects the two through the material and the project lineage rather than through shared terms. The traversal is visible while it runs, and each node in the answer resolves to its underlying record.

Expertise location works the same way. Rank the people who know about a topic and the ranking is assembled from evidence in the index: patents, experiments they led, publications, and, deliberately, their failed experiments. Add a second topic and the list re-orders rather than doubling, because evidence is deduplicated across topics. The person who can explain why a route failed is frequently not the person a skills spreadsheet would have named, since nobody lists a failed program under their areas of expertise.

Where a Person Stays in the Loop

A similarity alert flags prior work for the researcher to read. It does not block the plan. An 88% match can still be worth running, because the substrate differs, the target specification has moved, or the team wants to replicate a study with a single changed variable. In these cases, the researcher opens the earlier report and makes the call.

Introductions across sites are drafted and sent by a person. The draft is assembled from the topic and the expert's own evidence, the researcher edits it, and the request is logged with the topic that produced it.

Access holds at retrieval. Documents open under the reader's own permissions, and the same indexed fields a researcher filters on define which slice of the archive the R&D discovery agent may query. As a result, a researcher can’t surface an answer they are not authorized to read. We described the same pattern for governed retrieval in proposal automation, only here the sensitivity sits in unpublished results.

Four Searches to Run against Your Own Archive This Week

You can establish whether your organization has this problem without buying anything or running a formal audit. These four searches take a morning, and each one tests a different layer.

Run this

What it tests

A healthy result

What a poor result tells you

Search for a failure mode, not a project name: "thermal degradation," "phase separation," "poor adhesion"

Whether failure modes are indexed at all

Closed failed experiments return, ranked, with root cause visible

Your failures are stored, and they are unreachable by the only question that would find them

Search a material plus a process step, the way a researcher would describe a plan

Whether the archive answers questions about intended work

Results from more than one site, including work you did not know about

The archive answers by identifier rather than by intent, so plan checking is manual

Pick a researcher who left in the past two years and reconstruct what they concluded

Whether institutional knowledge survived their departure

Their reports, their failures, and whoever continued the work

Retention is a folder, and reading it requires someone who already knows the answer

Take a plan currently in review and ask what prior work it resembles

Whether prior-art checking happens before commitment

A ranked list with outcome, cost, and duration attached

The check is an opinion from whoever has been there longest

There’s one limit that deserves stating. Obviously, a plan check only sees what reached the index. Experiments that were never written up, or written up in a personal notebook that never left a laptop, stay invisible to it. What this changes is the return on a complete archive. It raises the value of every report your teams already write, rather than simply asking them to write more.

The rest is connection work: connectors into SharePoint and file shares, integration into ELN systems and patent sources, permission-aware retrieval throughout, deployable in your own cloud. Many organizations already have the records already, only they are sitting in three systems that have never been asked a question together.

See the R&D Knowledge Discovery Agent in a live demo, or read how a staged rollout works in practice in our finance transformation roadmap.

Frequently Asked Questions.

What is R&D knowledge management, and why does it fail to prevent duplicate research?
R&D knowledge management fails at duplication because archives are organized by what work was, rather than by what it found. A failure report is filed under an experiment ID and a material system, while its root cause sits in prose. Chalmers and Glasziou argued in The Lancet in 2009 that cumulative avoidable waste could reach 85% of research investment, partly through work that ignores prior findings.
How can an AI agent detect duplicate research before an experiment is funded?
The researcher describes the planned work in their own words, with materials and process steps, and the system scores that description against indexed metadata from every record in the archive. Matches above a similarity threshold return with their outcome, cost, and duration attached. Because scoring runs on indexed fields, the same plan returns the same score every time it is checked.
Do you need a knowledge graph or taxonomy before deploying an R&D knowledge discovery agent?
No. Grounded retrieval over an indexed archive carries the duplication check on its own, and that is the baseline for deployment. A knowledge graph adds reachability: it connects a question to prior work whose report never uses the same words, and it ranks experts from evidence rather than from a skills list. Organizations with an existing vocabulary get more sooner.
How do you make failed experiments searchable in an R&D archive?
Extract the failure mode, cost, duration, and recommendation at ingest and index them as fields alongside the document text. A failed-experiment report filed as a PDF is searchable only by its title and metadata, which name the project rather than the problem. Once the root cause is a field, a search for that failure mode returns the report that recorded it.
Who reviews what an R&D knowledge discovery agent finds, and is the process auditable?
A person decides at every point that carries consequence. A similarity alert flags prior work for a researcher to read, and they decide whether to proceed. Introductions to colleagues are drafted, edited, and sent by a person, with the request logged against its topic. Documents open under the reader's own permissions, and every aggregate figure resolves to its source records.
Can an R&D knowledge discovery agent connect to an ELN, SharePoint, or patent databases?
Yes. The agent runs over an index built from your existing systems, with connectors into SharePoint and file shares, integration into ELN systems, and ingestion of patent and publication sources. Most organizations already hold the records the check needs. They sit in three or four systems that have never been queried together, which is the gap the index closes.