通过 MCP 检索 MLflow 追踪并回读质量反馈

原文:https://mlflow.org/cookbook/mlflow-mcp-server/。作者/维护方:MLflow 官方网站(原页未署个人作者)。中文翻译整理与校注:未完纪;源文核对日期:2026-10-05。

MLflow 的 MCP Server 把追踪、实验和运行记录等操作变成 AI 助手可以调用的工具。启动 mlflow mcp run 后,客户端通过 MCP 的统一接口搜索失败请求、找出慢请求、读取 span,或给一条 trace 写入质量反馈。本文按官方 Cookbook 的完整流程展开,并保留其代码;其中涉及版本、身份与删除数据的边界,另作校注。

原文标题为 MLflow MCP Server: Debug, Analyze, and Annotate Traces from Any AI Assistant,页面日期为 2026-09-05。页面未提供可确认的个人作者署名,维护及发布方是 MLflow 官方网站。

AI 客户端通过 stdio 调用 mlflow mcp run 暴露的工具,服务端访问 MLflow 追踪与实验数据
图 1:MLflow 原文架构图。stdio 连接客户端与 MCP 子进程;工具集由 MLFLOW_MCP_TOOLS 控制。来源:MLflow 官方 Cookbook。图中标注 MLflow 3.15,正文最低版本声明为 3.5.1,两者不等于同一兼容性保证。

准备一个本地演示实验

原文给出的安装要求是:

pip install "mlflow[mcp]>=3.5.1"

接着在另一个终端启动 tracking server,并让它保持运行:

mlflow server --backend-store-uri sqlite:///mlflow.db --port 5000

这两条命令分别安装软件并启动服务,第二条会使用本地 mlflow.db 保存数据。本文没有执行它们。演示应使用独立目录与实验,避免与已有追踪数据混在一起。源码只写 >=3.5.1,没有固定版本;MCP Server 的最低引入版本不代表后文每个工具名、字段和分类都能在 3.5.1 使用。实际复现应固定一组兼容的 MLflow、FastMCP 与 Python 版本,并先检查 list_tools() 返回的 schema。

配置 Python SDK 和 MCP 子进程使用同一 tracking 地址,再选中名为 mcp-server-demo 的实验:

import os
import sys
import mlflow
from packaging.version import Version

TRACKING_URI = "http://localhost:5000"
os.environ["MLFLOW_TRACKING_URI"] = TRACKING_URI  # the MCP server subprocess reads this
mlflow.set_tracking_uri(TRACKING_URI)

experiment = mlflow.set_experiment("mcp-server-demo")

assert Version(mlflow.__version__) >= Version("3.5.1"), (
    f"The MLflow MCP Server needs MLflow >= 3.5.1 (found {mlflow.__version__})."
)
print(f"Experiment: {experiment.name} (id={experiment.experiment_id})")
# Example: Experiment: mcp-server-demo (id=1)

os.environ 中的 MLFLOW_TRACKING_URI 供稍后启动的子进程读取,mlflow.set_tracking_uri 则配置当前 Python 进程。set_experiment 可能复用同名实验,因此它不是“必定只包含本次数据”的保证。版本断言只检查下限;示例打印的实验 ID 也是示意值。

生成七条不调用模型的追踪

为了让工具有数据可查,原例使用普通的 @mlflow.trace 函数生成几种典型情况:正常的 LLM span、刻意等待约 1.3 秒的 TASK、主动抛出异常的 TOOL、包含检索和生成子 span 的 RAG CHAIN、带部署标签的追踪,以及手工记录 token 数的 CHAT_MODEL。这里的 LLM 名称只是 span 类型;函数返回固定字符串,不需要模型服务或 API key。

import time
from mlflow.entities import SpanType


@mlflow.trace(span_type=SpanType.LLM)
def answer(question: str) -> str:
    """A healthy (OK) trace tagged as an LLM span. We attach feedback to one later."""
    return f"Here is a helpful answer to: {question}"


@mlflow.trace(span_type=SpanType.TASK)
def slow_report() -> str:
    """An OK but deliberately slow (~1.3s) trace for the latency workflow."""
    time.sleep(1.3)
    return "generated a large report"


@mlflow.trace(span_type=SpanType.TOOL)
def flaky_tool() -> str:
    """Raises on purpose so MLflow records the trace with ERROR status."""
    raise RuntimeError("simulated downstream failure")


@mlflow.trace(span_type=SpanType.RETRIEVER)
def retrieve(question: str) -> list[str]:
    """A RETRIEVER child of the RAG trace below."""
    return ["doc: MLflow Tracing records spans", "doc: spans nest into a trace"]


@mlflow.trace(span_type=SpanType.LLM)
def generate(question: str, docs: list[str]) -> str:
    """An LLM child of the RAG trace."""
    return f"Based on {len(docs)} documents: here is the answer to '{question}'."


@mlflow.trace(span_type=SpanType.CHAIN)
def rag_answer(question: str) -> str:
    """A small RAG pipeline: a CHAIN parent with RETRIEVER and LLM children."""
    docs = retrieve(question)
    return generate(question, docs)


@mlflow.trace(span_type=SpanType.LLM)
def production_query(question: str) -> str:
    """Attaches custom tags so we can filter by deployment context."""
    mlflow.update_current_trace(tags={"environment": "production", "user_tier": "premium"})
    return f"Production answer to: {question}"


@mlflow.trace(span_type=SpanType.CHAT_MODEL)
def chat_with_usage(question: str) -> str:
    """Records token usage on its span; rolls up to trace.info.token_usage."""
    span = mlflow.get_current_active_span()
    span.set_attribute(
        "mlflow.chat.tokenUsage",
        {"input_tokens": 42, "output_tokens": 58, "total_tokens": 100},
    )
    return f"Chat answer to: {question}"

rag_answer 是根 span,内部依次调用 retrieve 和 generate。production_query 设置了 environment=production 与 user_tier=premium,这些只是合成演示标签,不表示代码真的请求了生产系统。chat_with_usage 设置输入 42、输出 58、总计 100 个 token,同样是手工写入的示例数值,不能当作实际模型计量。

随后运行这些函数,并在需要立刻检索或保留 ID 时刷新异步 trace 日志:

# A healthy trace we'll attach feedback to later.
answer("What is MLflow Tracing?")
mlflow.flush_trace_async_logging()
ok_trace_id = mlflow.get_last_active_trace_id()

# More OK / slow / error traces.
answer("How do I evaluate an agent?")
slow_report()
try:
    flaky_tool()
except RuntimeError:
    pass  # the failure is captured in the trace as an ERROR

# A nested RAG trace -- capture its id to inspect the span hierarchy later.
rag_answer("How does MLflow tracing work?")
mlflow.flush_trace_async_logging()
rag_trace_id = mlflow.get_last_active_trace_id()

# A tagged production trace and a chat trace with token usage.
production_query("What is my order status?")
chat_with_usage("Summarize the MCP docs.")
mlflow.flush_trace_async_logging()
chat_trace_id = mlflow.get_last_active_trace_id()

print(f"Seeded 7 traces in experiment {experiment.experiment_id}")
# Example: Seeded 7 traces in experiment 1

这里产生七条根 trace:两次正常回答、一次慢任务、一次失败工具、一次 RAG、一次带标签的请求和一次带 token 属性的聊天。flaky_tool 的异常被捕获,MLflow 仍会把错误状态写入追踪。代码保留 ok_trace_id、rag_trace_id 和 chat_trace_id,供后文精确读取。注释中的 “Seeded 7 traces” 来自上游演示,不是本文的运行证明。

通过 stdio 连接 MCP,并先列出工具

mlflow mcp run 使用 stdio:客户端启动一个子进程,通过标准输入输出管道交换 JSON-RPC 消息。这个传输本身不需要为 MCP 新开监听端口,子进程通常随客户端会话结束。但 MCP 进程仍然需要访问 tracking server;如果 tracking 地址是远程服务,它照样会发生网络通信。

MLFLOW_MCP_TOOLS 控制工具类别。减少工具数量可以降低助手接收工具定义的 token 开销。以下数量是原文版本中的示例,实际以 list_tools() 为准:

配置值 类别 原文数量
genai(默认) traces、scorers、experiments、runs 26
ml experiments、runs、models、deployments 32
all 上述全部 45
逗号分隔的类别 例如 traces,scorers 随选择变化

原例使用 FastMCP,定义每次打开客户端、调用一个工具并取回文本的辅助函数。Notebook 或已有事件循环的 REPL 使用 nest_asyncio;普通 .py 脚本则应按注释去掉该导入及 apply()。这些辅助包也需要存在于所选环境中。

import asyncio
import nest_asyncio
from fastmcp import Client
from fastmcp.client.transports import StdioTransport

# In a notebook or REPL an event loop is already running; nest_asyncio lets us
# call asyncio.run() inside it. Omit this line (and the import) in a plain .py script.
nest_asyncio.apply()


def mcp_transport(tools: str = "genai") -> StdioTransport:
    """Launch `mlflow mcp run` over stdio with a chosen MLFLOW_MCP_TOOLS set."""
    return StdioTransport(
        command=sys.executable,  # `python -m mlflow`; no dependency on `mlflow` being on PATH
        args=["-m", "mlflow", "mcp", "run"],
        env={**os.environ, "MLFLOW_MCP_TOOLS": tools},
    )


def mcp(tool: str, arguments: dict | None = None, tools: str = "genai") -> str:
    """Call one MLflow MCP tool and return its text output."""
    async def _call():
        async with Client(mcp_transport(tools)) as client:
            result = await client.call_tool(tool, arguments or {})
            return result.content[0].text

    return asyncio.run(_call())


async def _list_tools():
    async with Client(mcp_transport("genai")) as client:
        return await client.list_tools()


tools = asyncio.run(_list_tools())
print(f"{len(tools)} tools exposed with MLFLOW_MCP_TOOLS=genai")
# Example: 26 tools exposed with MLFLOW_MCP_TOOLS=genai
# search_traces, get_trace, delete_traces, set_trace_tag, log_trace_feedback,
# get_trace_assessment, evaluate_traces, list_scorers, create_experiment,
# search_experiments, list_runs, create_run, link_traces_to_run, ...

辅助函数使用当前 sys.executable 启动 python -m mlflow,避免 PATH 指向别的 MLflow。它把父进程环境传给子进程,因此不应在不可信客户端下继承无关敏感变量。代码还假设返回内容的第一项一定有 text;健壮客户端应先检查工具错误和内容类型,再处理所有需要的内容块。它每次调用都会建立新会话,适合教学,但并非高吞吐会话管理方案。

权限边界:genai 和 traces 不是只读权限。原列表就包含 delete_traces、log_trace_feedback、创建实验与运行等写操作。工具类别限制不能替代 tracking 服务端权限控制。应使用最小权限凭据、限定实验,并让客户端对写入或删除操作进行明确审批。

查找失败请求

助手可以把“找出这个实验中的失败 trace”转为一次 search_traces 调用。filter_string 使用 MLflow 追踪搜索语法,extract_fields 限制返回的字段,避免把整条追踪都放进响应。

print(mcp("search_traces", {
    "experiment_id": experiment.experiment_id,
    "filter_string": "status = 'ERROR'",
    "extract_fields": "info.trace_id,info.state,info.request_preview",
}))
# Example:
# info.trace_id                        info.state    info.request_preview
# -----------------------------------  ------------  --------------------
# tr-f3912f13ffa051dd0cab558a42b32989  ERROR         {}

过滤条件中的 status 与输出列中的 info.state 处在不同接口位置,不应因名字不一样而自行替换。上游示例返回一个 ERROR trace,并用精简表格显示 ID、状态和请求预览。

查找执行超过一秒的追踪

慢请求使用毫秒阈值过滤。返回列名为 info.execution_duration_ms,原文展示的格式化输出是 1.3s:

print(mcp("search_traces", {
    "experiment_id": experiment.experiment_id,
    "filter_string": "execution_time_ms > 1000",
    "extract_fields": "info.trace_id,info.execution_duration_ms",
}))
# Example:
# info.trace_id                        info.execution_duration_ms
# -----------------------------------  --------------------------
# tr-0cc558dc15db559253da5d9b4917f8ce  1.3s

这个结果对应刻意 sleep(1.3) 的模拟任务,不是一次真实模型或数据库性能测试。将同样的查询用到实际实验时,还需要确认时间范围、采样策略和请求类型。

写入质量反馈,然后回读确认

log_trace_feedback 为 trace 增加一条 assessment,可用于记录人工评审或模型评审结果。该工具的 value 以 JSON 编码字符串传入,所以数值 0.9 写成字符串 "0.9"。写入后,再用 get_trace 只取 assessments 字段,检查保存的名称、来源、分值和理由。

print(mcp("log_trace_feedback", {
    "trace_id": ok_trace_id,
    "name": "relevance",
    "value": "0.9",              # JSON-encoded value
    "source_type": "HUMAN",
    "source_id": "reviewer@example.com",
    "rationale": "Answer is on-topic and accurate.",
}))
# Example: Logged feedback 'relevance' to trace tr-a3a3... Assessment ID: a-f2f8...

print(mcp("get_trace", {
    "trace_id": ok_trace_id,
    "extract_fields": "info.assessments.*",
}))
# Example (abridged):
# {"info": {"assessments": [{"assessment_name": "relevance",
#   "source": {"source_type": "HUMAN", "source_id": "reviewer@example.com"},
#   "feedback": {"value": 0.9}, "rationale": "Answer is on-topic and accurate."}]}}

身份校注:上面的 source_type="HUMAN"、reviewer@example.com 和“答案切题且准确”的理由均是原文占位演示。没有真实人工评审时,不能用这组字段伪装已经有人审阅;应使用实际来源及版本支持的来源类型,并明确这是合成数据。回读只能确认记录了什么,不能证明评价内容正确。

按标签定位、展开 span 层级、读取 token 属性

默认 genai 工具集还包括实验和运行管理。下面列出最多五个实验:

print(mcp("search_experiments", {"max_results": 5}))

再按环境标签找出前面人工标记为 production 的追踪:

print(mcp("search_traces", {
    "experiment_id": experiment.experiment_id,
    "filter_string": "tags.environment = 'production'",
    "extract_fields": "info.trace_id,info.tags.environment,info.tags.user_tier",
}))
# Example:
# info.trace_id                        info.tags.environment    info.tags.user_tier
# -----------------------------------  -----------------------  -------------------
# tr-e7415d0d5886aeab90c82dc0b13558d7  production               premium

RAG 的根 span 没有父级,RETRIEVER 与 LLM 子 span 的 parent_span_id 指回根节点。读取名称、span ID 和父 ID,就能还原调用关系:

print(mcp("get_trace", {
    "trace_id": rag_trace_id,
    "extract_fields": "data.spans.*.name,data.spans.*.span_id,data.spans.*.parent_span_id",
}))
# Example (abridged): rag_answer (root) -> retrieve, generate (both parented to rag_answer)

token 用量保存在 span 的标准属性 mlflow.chat.tokenUsage 中。原例说明字段选择器可以返回完整 attributes 映射,但不能直接用带点的属性键完成这一层精确选择;因此先读取属性,再对其中的 JSON 编码字符串解码。

import json

detail = mcp("get_trace", {
    "trace_id": chat_trace_id,
    "extract_fields": "data.spans.*.attributes",
})

for span in json.loads(detail)["data"]["spans"]:
    raw_usage = span.get("attributes", {}).get("mlflow.chat.tokenUsage")
    if raw_usage:
        print(json.loads(raw_usage))
        # Example: {'input_tokens': 42, 'output_tokens': 58, 'total_tokens': 100}

这段读取逻辑依赖原文工具版本返回的编码形式。如果实际 schema 或返回对象已经给出字典,应按实际类型处理,避免重复 JSON 解码。在 Python SDK 中,mlflow.get_trace(id).info.token_usage 是更直接的访问方式;MCP 工具则让只有外部工具接口的助手也能访问同一类数据。

注册到 AI 客户端

实际使用时通常不必手写客户端。注册服务后,助手可以根据自然语言调用工具。MLFLOW_TRACKING_URI 可以指向本地服务器、远程 URL,或 Databricks。Databricks 场景还需配置 DATABRICKS_HOST 和 DATABRICKS_TOKEN,但不要把真实 token 写进会提交到项目仓库的 MCP 配置。

原文给出的命令使用 claude CLI:

claude mcp add mlflow-mcp -e MLFLOW_TRACKING_URI=http://localhost:5000 \
  -- uv run --with "mlflow[mcp]>=3.5.1" mlflow mcp run

该 CLI 示例适用于有相应命令支持的 Claude Code 环境;原文将 Code/Desktop 放在同一标题下,不应据此断言 Desktop 的所有版本都接受这一 CLI 配置方式。下面则是原文项目级 .mcp.json 的配置结构:

{
  "mcpServers": {
    "mlflow-mcp": {
      "command": "uv",
      "args": ["run", "--with", "mlflow[mcp]>=3.5.1", "mlflow", "mcp", "run"],
      "env": {
        "MLFLOW_TRACKING_URI": "http://localhost:5000",
        "MLFLOW_MCP_TOOLS": "genai"
      }
    }
  }
}

原文提到 VS Code 使用 servers 顶层键,Cursor 使用 .cursor/mcp.json。具体客户端版本可能还要求不同位置、类型字段或信任确认,应按该客户端当前文档核对,不要只机械复制顶层键。这里的 uv run --with 还可能在启动时解析或下载依赖,演示中的最低版本约束同样需要替换为经过确认的版本锁定策略。

注册成功后,可以请求“找出实验 1 最近一小时的失败追踪”“列出执行超过一秒的最慢请求”,或“为指定 trace 记录 0.85 的相关性分数并写明理由”。时间范围、排序和写入对象应落实到真实工具参数;一句自然语言不能替代对调用范围的核对。连接外部 AI 助手、远程 tracking 或调用评价工具,还可能涉及模型费用与追踪内容外发;无 API key 的结论只适用于前面的合成种子函数。

原文的清理操作:默认不执行

delete_traces 可以按 trace ID 或最大时间戳删除数据。原文为了重置演示环境,使用“当前实验中,截至现在的全部追踪”作为删除范围:

原文破坏性清理示例,仅供审阅,不属于建议执行步骤
import time

print(mcp("delete_traces", {
    "experiment_id": experiment.experiment_id,
    "max_timestamp_millis": int(time.time() * 1000),
}))
# Example: Deleted 7 trace(s) from experiment 1.

如果 mcp-server-demo 已经存在,这条命令会包括其中早于本次演示的追踪,不能保证只删除刚生成的七条。本文将它从默认执行流程移除,保留在折叠的原例中以完整交代源文。需要清理时,应先备份,再核对本次实际生成的 ID 和当前工具 schema,按这些 ID 精确删除。原例只保留了三条 ID,不足以构成全部七条的安全清理清单;本次也未执行任何删除。

这套流程能确认什么

按原教程复现后,可以用同一套 MCP 接口检索失败与慢请求、查看标签和 span 关系、读取 token 字段,并写入一条反馈后回读确认。工具集选择减少暴露给助手的工具定义,字段选择缩小单次响应;真正的数据访问范围和写权限仍要由服务端与客户端策略控制。

进一步阅读原文链接的 MLflow MCP Server 文档、Search Traces Reference,以及 MLflow Tracing 的生产可观测性指南。本文是完整中文整理与静态技术审查,没有安装依赖、启动服务、连接客户端、生成 trace、写入反馈或删除数据。代码中的结果注释全部为上游示例。

来源与权利说明

来源:MLflow 官方 Cookbook 原文;© 2025 MLflow Project, a Series of LF Projects, LLC。原页未给出可确认的个人作者,也未见独立的文章转载许可文本,因此不把 MLflow 软件仓库的开源许可自动解释为文章与配图许可。源图原样保存并保留归属;没有伪造客户端界面或运行截图。

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容