CVE-2026-61539 Security Alert: CRITICAL Vulnerability

Urgent: CVE-2026-61539 requires immediate attention.

· 7 min read

Executive Summary

CVE-2026-61539 is a critical remote code execution flaw in Xinference affecting its OpenAI-compatible chat API path when parsing Llama3 tool-call output. In vulnerable deployments, Xinference used Python’s unsafe eval() on model-generated text. An attacker who can send crafted prompts to /v1/chat/completions may be able to make the server execute arbitrary Python expressions as the Xinference process user.

This is especially urgent for solo developers and small teams because the tested default deployment reportedly had no authentication enabled, making the issue reachable over the network without credentials. Even though there is no known exploitation in the wild and it is not in CISA KEV, the CVSS 10.0 score means you should treat this as an emergency patch-and-contain event.

Immediate Action

  • Upgrade Xinference immediately to a fixed release as soon as one is available. If no patched version is published yet, disable or isolate the exposed chat API until you can update. TODO: vendor advisory / fixed version link
  • Remove public access to /v1/chat/completions at the firewall, reverse proxy, or security group. Limit access to trusted IPs only.
  • Enable authentication if your deployment supports it, and place the service behind a private network or VPN.
  • Assume compromise is possible if the endpoint was internet-facing. Review logs, rotate secrets, and inspect the host for unexpected processes or outbound connections.
  • Rollback only if needed to a known-safe build that does not use the vulnerable parser path; otherwise prefer patching over rollback to avoid reintroducing older issues.
  • Alert your team and add temporary monitoring for unusual chat prompts, shell activity, and file changes in the Xinference runtime environment.

Affected Versions

  • xinference@<TODO-fixed-version vulnerable; upgrade to TODO-fixed-version or later.
  • pip xinference installations using the affected Llama3 tool-call parser are vulnerable when the chat API is reachable.
  • Safe versions: TODO-fixed-version+ with the eval() usage removed or replaced by safe parsing.

Resolution Guide

Python / pip

python -m pip install --upgrade xinference
python -m pip show xinference

pipx

pipx upgrade xinference
pipx list

Docker

docker pull xinference/xinference:TODO-fixed-tag
docker stop xinference
docker rm xinference
docker run --restart unless-stopped -p 9997:9997 xinference/xinference:TODO-fixed-tag

Linux package managers (if your environment repackages Xinference)

sudo apt-get update
sudo apt-get install --only-upgrade xinference

sudo yum update xinference

JavaScript package managers are not typically used for Xinference itself, but if you have a wrapper service or client, update it separately:

npm update
yarn upgrade
pnpm update

Hardening steps

# Put Xinference behind a reverse proxy with auth and IP allowlisting
# Example Nginx snippet:
location /v1/chat/completions {
    allow 10.0.0.0/8;
    deny all;
    proxy_pass http://127.0.0.1:9997;
}
# If your deployment supports config flags, disable tool calling or the affected parser path
# TODO: replace with the exact Xinference setting for your version
export XINFERENCE_DISABLE_TOOLS=true
export XINFERENCE_REQUIRE_AUTH=true

Minimal code fix example — replace eval() with safe parsing:

import ast
from typing import Any, Dict, List, Optional, Tuple

def extract_tool_calls(self, model_output: str) -> List[Tuple[Optional[str], Optional[str], Optional[Dict[str, Any]]]]:
    try:
        data = ast.literal_eval(model_output)
        if not isinstance(data, dict):
            raise ValueError("tool call must be a dict")
        return [(None, data["name"], data["parameters"])]
    except Exception:
        return [(model_output, None, None)]

Detection & Verification

Check installed version

python -m pip show xinference
python -c "import xinference; print(xinference.__version__)"

Look for the vulnerable pattern in source or site-packages

grep -R "eval(model_output" -n /usr/local/lib/python*/site-packages/xinference 2>/dev/null
grep -R "extract_tool_calls" -n /usr/local/lib/python*/site-packages/xinference 2>/dev/null

Verify the fix

# Confirm the vulnerable eval() call is gone
grep -R "eval(model_output" -n /usr/local/lib/python*/site-packages/xinference 2>/dev/null

# Confirm the service no longer accepts unauthenticated public requests
curl -i http://YOUR_HOST:9997/v1/chat/completions

Dependency audit

python -m pip install pip-audit
pip-audit | grep -i xinference

Runtime checks

# Look for suspicious child processes or outbound connections from the Xinference PID
ps auxf | grep xinference
ss -plant
lsof -p $(pgrep -f xinference | head -n1)

Risk and Impact

If exploited, this bug can give an attacker arbitrary code execution inside the Xinference server process. That can lead to theft of API keys, model files, prompt data, environment secrets, and any other data the service can read. Because the issue is network-reachable and requires no user interaction, a single malicious chat request may be enough to compromise the host.

For small teams, the blast radius can include the entire development environment if Xinference runs with broad filesystem access, shared credentials, or cloud metadata permissions. Treat any exposed instance as high risk until patched and verified.

Keep reading