![]() |
| Beware the MCP trojan horse |
As patent attorneys, we should all know by now not to put client confidential information to a non-enterprise version of an LLM. However, as the capabilities of AI tools become more complex, so too does the necessity of understanding what exactly they are doing and where our data goes, beyond the simple prompt. One such requirement is understanding the process by which AI tools silently retrieve information from the web and external databases in response to user queries via MCPs. Such processes often fall outside of the confidentiality provisions governed by the LLM per se. We therefore need to be asking our AI provider of choice what information is sent out to external parties, including web search engines and the confidentiality provisions associated with this. After all, web search is web search, whether this is a Google search, instructed by an LLM in response to a user query, or carried out by legal tech software sending instructions to a third-party tool.
LLMs and information retrieval
A large language model (LLM) is, at bottom, a text-in, text-out machine. Its knowledge is frozen at the point its training data stopped, it has no memory of your file, and in its raw state no way of looking anything up. Left to itself, asked about the prosecution history of a European application, it will produce something that reads exactly like a prosecution history and is very often fiction (IPKat). An MCP is an agreed way of plugging an AI model into something else, whether a database, a search engine, a website or your own document system, so that the model can go and fetch what it needs rather than relying on whatever it happens to remember from its training.
An MCP server is a single connector, built either by whoever holds the data or by a third party wrapping a public API to that data. LLMs can engage with this connector, which sets out a menu of things the model may ask for, such as searching publications, fetching register information or downloading the file wrapper.
Readers who have used Claude, ChatGPT or Gemini in the last year will object that the models plainly do look things up. This is because most state-of-the-art LLMs now have web-search built in. In response to a query, they will conduct a web search and cite their sources. This built-in web search is an example of the mechanism operating behind MCPs. The model is not remembering or reading the internet. It is interpreting your query, calling a web search application, getting text back, and reasoning over whatever it received. The request sent by the LLM to the web search application, known as a tool call, is a short, structured instruction the model composes for itself, naming the tool it wants and the arguments to run it with. The web search application then executes it on the model's behalf and feeds the answer back. This mechanism operates in the same way whatever the tool or application happens to be.
MCPs for patent work
The patent profession is unusually well served by public data. It is the name of the game that patents are published, and these publications are freely available for anyone to read from their open access database of choice. Patent prosecution for published cases may also be freely viewed and downloaded from patent office databases. As patent attorneys, we also need to be familiar with the databases of technical information relevant to our field, such as databases of sequences, chemical structures, clinical trials, drug approvals and regulatory guidance.
However, the information provided in most of these databases is difficult to access, either for providing to models during a training run or by connecting the models to a simple web search tool. This is why even the latest LLMs with access to web search struggle to reliably and accurately retrieve patent data. What a model can reach is only ever as good as the tools it has been handed, and a general web search is a blunt instrument for patent purposes. Much early disappointment with LLMs in patent work has really been disappointment with retrieval. Enter the MCP servers. MCP connectors to patent and scientific databases allow LLMs to access and accurately provide patent information in a way that is not possible either by model training or simple web search.
There are already many MCP servers available for patent work, including free connectors to the EPO's Open Patent Services (OPS) API. Note, however, that whilst the data and the underlying API are provided by the patent office, the MCP server wrapping them is generally built and operated by someone else. You can also get MCP servers to connect your LLM of choice to scientific databases such as, for example, Pubmed, PubChem, NCBI BLAST sequence search and arXiv. Commercial patent intelligence databases and AI wrapper companies are also now offering their own MCP servers for LLMs.
So what’s the catch?
If you are using an LLM such as ChatGPT, Claude or Gemini for professional patent work you should be using an enterprise account to maintain the confidentiality of your client’s data. Information submitted to a free-to-use non-enterprise model is generally considered a non-confidential disclosure (IPKat). If you are using an AI legal tech / AI wrapper company for professional patent work, your company of choice should have the necessary agreements in place with their LLM providers to preserve the confidentiality of users’ inputs. Of course, you also need an agreement with the AI wrapper company that they themselves will protect your data.
So how do MCPs fit into this? Looking first at the LLM providers themselves. If you connect your LLM to a third-party MCP server, this generally is not covered by the terms and conditions that govern your use of the model itself. This is essentially because the MCP provider is a third party in the relationship. For example, if you connect Claude to a third-party MCP server for patent search, you lose control of the data that Claude sends to the MCP server. As Anthropic states themselves: “Anthropic does not review, authorize, or assume responsibility for customer-registered MCP servers or the external systems they call”.
The information sent by the LLM to the MCP server, in the case of patent search, can contain very sensitive information, such as a description of the undisclosed invention, a proprietary sequence, a new disease target or other confidential client information. Connecting even an enterprise-level LLM account to a third-party MCP server is therefore effectively like submitting the details of the invention into a Google search or a free online database. Where the server is hosted by someone else, whoever operates it receives your tool calls and may log them. In a simpler example, the same is true if you enable web search in Gemini or Claude. The information is sent out to the third-party web search provider and may be logged by them in the manner of all web searches.
Why do we care?
In some respects, this problem is not new. As patent attorneys we know that it is generally best to avoid putting the specifics of an unfiled, novel invention into a public search engine. Whilst typing terms into a search engine does not typically constitute a "public disclosure", it is generally considered to violate best practices for maintaining absolute confidentiality.
The first problem is data logging. Public search engines log queries, IP addresses, and user profiles. If the target involves a highly specific, newly identified protein name, gene sequence, or structural formulation, querying it directly transmits that data to a third-party server without a non-disclosure agreement.
The second problem is trade secret vulnerability. If you are holding the target as a trade secret whilst developing the assays or drug candidates, exposing the exact nomenclature or sequence to a public platform weakens the argument that reasonable steps were taken to keep the information secret. This is why you don’t submit a BLAST search for your newly invented antibody sequence.
Connecting to a third-party MCP server for database searches without specific confidentiality agreements with the MCP provider falls under exactly the same category.
The risks are real
So you are using a legal tech / IP AI wrapper company for patent work. You know they have confidentiality and no-data-retention agreements in place with their LLM providers. You also have a specific confidentiality agreement with the legal tech company themselves. So far so good.
However, most new legal tech companies will not have their own database of curated IP and scientific resources. They will therefore be using some form of MCP server to access patent and scientific databases. The next questions you need to be asking are therefore how they are accessing the data they are providing:
Establishing this information is particularly important if you are using the tool to carry out invention disclosure assessments.
There are currently IP AI-wrapper companies which offer MCP connectors to their system from Claude (you generally need a subscription to both). The connectors say that they allows users to search across patent and non-patent literature, legal texts and case law, SEP technical standards, "and the open web”. However, this Kat thinks it is unlikely that legal tech companies have built their own web search engines, or a comprehensive database of all the journal articles, theses, datasets and books. It seems likely, therefore, that MCP servers are in play. In order to be confident in submitting confidential client information such as an invention disclosure to these connectors, it is necessary to confirm that the data is not sent to third-party or web search engines, unless you would be similarly happy to submit the same information in a Google search.
Final thoughts
According to the epi Guidelines on the Use of Generative AI in the Work of Patent Attorneys, “Members must inform themselves about the likelihoods and modes of non-confidential disclosures deriving from use of specific AI models”. In the ever-changing AI landscape, this can be easier said than done. MCPs add yet more complexity for patent firms to navigate. The Trojan horse is that the connector arrives inside a carefully negotiated confidential arrangement and quietly carries the data back out again. Ironically, negotiating and putting in place a whole load of patent-related MCP servers with the required confidentiality provisions in place is one of the few things legal tech companies have to offer over and above the LLM providers themselves (IPKat). Commercial databases of curated patent information may find that their value will lie almost exclusively in how well they enable users to connect to their databases from LLMs and how stringent their associated confidentiality provisions are. Perhaps someone could just offer this, minus the cost of a full subscription to wrapper functionality we no longer need?
Acknowledgements: Thanks, as always, to Dr Laurence Aitchison (Head of Reasoning at Mistral) for explaining AI.
Further reading
Use of AI in the patent industry series: