Unpatched LMCache Flaw Allows Unauthenticated Remote Code Execution

JFrog disclosed CVE-2026-105192, a CVSS 9.8 pickle deserialization flaw in LMCache that lets one network message run code. No patched version exists yet.
Table of Contents
    Add a header to begin generating the table of contents

    JFrog Security Research has disclosed a critical, unpatched vulnerability in LMCache, a caching layer used with large language model inference, that lets an unauthenticated attacker run code on a server with a single crafted network message.

    The flaw is tracked as CVE-2026-105192 and carries a CVSS score of 9.8. JFrog published its findings on October 7, and no patched version of LMCache exists.

    CVE-2026-105192 Abuses Pickle Deserialization Over ZeroMQ

    According to JFrog, the vulnerability comes from unsafe deserialization of Python pickle data in LMCache’s multiprocess mode, which communicates over ZeroMQ. A pickle payload can carry instructions that run when the data is loaded, so a server that deserializes untrusted pickle data will execute whatever an attacker places in it.

    Yuval Moravchick of JFrog Security Research found the flaw. A single crafted message is enough, and the attacker needs no authentication. The code runs with the privileges of the LMCache process, which are root privileges in the project’s official containers.

    Which LMCache Versions Are Affected

    The flaw affects LMCache versions 0.3.9 through 0.5.5. It also affects the 0.5.6 release candidates and the development branch. No fixed release has been published.

    The Condition That Limits Exposure

    The vulnerability can be exploited only when the LMCache server listens on a routable network address instead of localhost. Deployments bound to localhost are not reachable by the network attack JFrog describes. The practical risk therefore falls on clusters where the multiprocess-mode port is open to networks that include untrusted hosts.

    The attack surface is narrow but consequential. LMCache’s multiprocess mode exchanges messages over ZeroMQ, and the flaw exists in how those messages are processed. An attacker who can reach the listening port does not need a login, a token or any prior foothold: the crafted message alone triggers the code execution. Because the official containers run the process as root, the attacker’s code inherits that level of privilege from the first moment it runs.

    More Reports Remain Unconfirmed

    Six additional reports were filed on October 6 covering unauthenticated cache access and services that execute commands. According to the disclosure, the LMCache maintainers have not confirmed those reports. The disclosure does not say how the maintainers responded to the CVE-2026-105192 report itself.

    The reports leave open the possibility that the project has further security problems beyond the deserialization flaw. Without confirmation from the maintainers, the scope of those reports cannot be assessed from the public record.

    Impact on AI Inference Clusters

    JFrog’s assessment is that AI inference clusters exposing the LMCache port to untrusted networks are open to full takeover. Because the process runs as root in official containers, code execution translates directly into control of the container and whatever the container can reach.

    Inference infrastructure often runs on expensive GPU hardware and handles model data and prompts, so a compromise carries both compute and data implications. The disclosure as summarized does not mention any attacks in the wild.

    JFrog’s Mitigation Guidance

    With no patch available, JFrog advises binding LMCache to localhost or to trusted cluster networks and firewalling the port. These steps remove the network path the attack requires, which is the one condition on which the exploit depends.

    Pickle deserialization flaws have a long history across Python software, and the pickle format’s documentation warns against loading data from untrusted sources. The LMCache case applies that known risk to a network-facing component in AI serving stacks, where components are often wired together across hosts in a cluster.

    Operators running LMCache in multiprocess mode can check the version against the affected range and confirm what network interface the service listens on. Until the maintainers release a fix, restricting network reach is the only mitigation JFrog describes.

    Related Posts