用 DSPy 从仓库元信息生成 llms.txt 草稿

llms.txt 可为模型或检索工具提供项目导览,说明用途、关键概念、架构和示例入口。DSPy 教程把仓库信息转换成结构化输入,再用签名与模块组织分析。生成结果应是待审草稿,不是权威 API 文档。

安全读取仓库元数据后由 DSPy 生成 llms.txt 草稿,再由维护者人工审阅。
未完纪原创流程图。

拆分分析任务

AnalyzeRepository 接收仓库 URL、文件树和 README,输出用途、关键概念和架构概述;AnalyzeCodeStructure 从文件树与包文件总结目录、入口和开发信息;GenerateLLMsTxt 汇总分析与用法示例,生成文档正文。RepositoryAnalyzer 串接 ChainOfThought 预测器。源页最终示例按 Project Overview、Key Concepts、Architecture、Usage Examples 组织。

仓库采集示例解析 owner/repo,通过 GitHub REST API 读取 tree、README 和常见包文件,再把内容交给 DSPy。真实服务应只允许 GitHub API 主机、限制仓库路径与读取文件范围,并核对 GitHub contents API 对大文件的不同响应形态。

修复凭据和请求处理

原文用 os.environ["GITHUB_ACCESS_TOKEN"] = "<your_access_token>" 与 os.environ["OPENAI_API_KEY"] = "<YOUR OPENAI KEY>" 设置占位密钥。这会覆盖已有凭据;把真实值直接写进源码又容易泄露。应从密钥管理器或运行环境读取,缺少时失败,并避免打印密钥。

token = os.environ.get("GITHUB_ACCESS_TOKEN")
if not token: raise RuntimeError("GITHUB_ACCESS_TOKEN is required")
response = requests.get(api_url, headers={"Authorization": f"Bearer {token}"}, timeout=(3.05, 20))
response.raise_for_status()
payload = response.json()

上述片段是本稿的安全修订,不是原文原样代码,也未经执行。源示例无统一 timeout/status 检查且有裸 except,可能隐藏超时、限流、JSON 错误和程序缺陷。应捕捉预期异常、限制响应大小、验证字段/编码/base64 并用有限退避。原始函数始终构造 api.github.com 地址,不能据此认定原代码已存在任意目标 SSRF;若以后把它扩展成接受任意抓取地址的服务,才需额外防范 SSRF。

把模型输出当草稿

README、文件名和代码都属于不可信输入。提示词应明确其内容是数据,不是要执行的指令;生成后人工核实事实、版本、链接和示例,并先写临时文件供 diff,避免覆盖仓库文档。源页使用当前 DSPy 文档路径、gpt-4o-mini 和当前 API,均会变化;复现时锁定依赖并核对 SDK。未执行教程中的 GitHub API 或模型调用。

来源:DSPy 文档,Generating llms.txt。静态审查列出了原文差异和安全修改。

本文补充:核对与维护细节

签名输入输出与生成顺序

原教程的 AnalyzeRepository 输入 repo_url、file_tree、readme_content,输出 project_purpose、key_concepts、architecture_overview。AnalyzeCodeStructure 输入文件树与 package files,输出目录、entry points、development info。生成用法示例模块先接收用途与概念,再把用途、概念、架构和 usage_examples 交给最终生成模块。由于模型输出不确定,应先保存中间 Prediction 用于审查,而不是只留最后字符串。

教程逐步获取 repository tree、README 与包清单,并对文件内容做 base64 解码;真实 GitHub 请求需尊重 API 限速与 404/403 语义,匿名和已认证配额不同。只收集完成生成真正需要的文件,避免把整个仓库、二进制文件、生成物或用户数据送入外部模型。生成 llms.txt 前还应让维护者确认仓库许可证、文档引用与未公开代码是否可以外发。

文件写入边界

原文将生成文本写到 llms.txt 并展示开头部分。自动化版本可先写入同目录临时文件,校验大小、编码和目标仓库根路径后生成 diff,审批通过再替换;若目标已存在,默认不要静默覆盖。模型回复截断或返回空内容时应显式失败,不要把部分文档当完整文件。

签名、DSPy 模块与采集流程

教程把分析拆成有明确 I/O 字段的签名,再由 RepositoryAnalyzer.forward 串起 ChainOfThought 预测器。核心数据契约如下;模型输出会随 DSPy/模型版本、提示词和仓库内容变化:

import dspy
from typing import List

class AnalyzeRepository(dspy.Signature):
    """Analyze a repository structure and identify key components."""
    repo_url: str = dspy.InputField(desc="GitHub repository URL")
    file_tree: str = dspy.InputField(desc="Repository file structure")
    readme_content: str = dspy.InputField(desc="README.md content")
    project_purpose: str = dspy.OutputField(desc="Main purpose and goals")
    key_concepts: list[str] = dspy.OutputField(desc="Important concepts")
    architecture_overview: str = dspy.OutputField(desc="High-level architecture")

class AnalyzeCodeStructure(dspy.Signature):
    """Analyze important directories and files."""
    file_tree: str = dspy.InputField(desc="Repository file structure")
    package_files: str = dspy.InputField(desc="Package/configuration files")
    important_directories: list[str] = dspy.OutputField(desc="Key directories")
    entry_points: list[str] = dspy.OutputField(desc="Main entry points")
    development_info: str = dspy.OutputField(desc="Development workflow")

class GenerateLLMsTxt(dspy.Signature):
    """Generate comprehensive llms.txt documentation."""
    project_purpose: str = dspy.InputField()
    key_concepts: list[str] = dspy.InputField()
    architecture_overview: str = dspy.InputField()
    important_directories: list[str] = dspy.InputField()
    entry_points: list[str] = dspy.InputField()
    development_info: str = dspy.InputField()
    usage_examples: str = dspy.InputField(desc="Usage patterns and examples")
    llms_txt_content: str = dspy.OutputField(desc="Complete llms.txt content")

class RepositoryAnalyzer(dspy.Module):
    def __init__(self):
        super().__init__()
        self.analyze_repo = dspy.ChainOfThought(AnalyzeRepository)
        self.analyze_structure = dspy.ChainOfThought(AnalyzeCodeStructure)
        self.generate_examples = dspy.ChainOfThought(
            "repo_info -> usage_examples")
        self.generate_llms_txt = dspy.ChainOfThought(GenerateLLMsTxt)

    def forward(self, repo_url, file_tree, readme_content, package_files):
        repo_analysis = self.analyze_repo(
            repo_url=repo_url, file_tree=file_tree,
            readme_content=readme_content)
        structure_analysis = self.analyze_structure(
            file_tree=file_tree, package_files=package_files)
        usage_examples = self.generate_examples(
            repo_info=(f"Purpose: {repo_analysis.project_purpose}\n"
                       f"Concepts: {repo_analysis.key_concepts}"))
        llms_txt = self.generate_llms_txt(
            project_purpose=repo_analysis.project_purpose,
            key_concepts=repo_analysis.key_concepts,
            architecture_overview=repo_analysis.architecture_overview,
            important_directories=structure_analysis.important_directories,
            entry_points=structure_analysis.entry_points,
            development_info=structure_analysis.development_info,
            usage_examples=usage_examples.usage_examples)
        return dspy.Prediction(
            llms_txt_content=llms_txt.llms_txt_content,
            analysis=repo_analysis, structure=structure_analysis)

采集阶段读取 GitHub tree、README 与 pyproject.toml、setup.py、requirements.txt、package.json;入口创建 dspy.LM(model="gpt-4o-mini")、配置 DSPy、调用采集与分析后写出 llms.txt。原文将 GITHUB_ACCESS_TOKEN 和 OPENAI_API_KEY 的占位值直接赋给环境变量,会覆盖已有密钥;如把真实值写进源码则硬编码泄露。源代码对 HTTP 未统一设 timeout/status,并使用裸 except,可能把失败伪装成缺失文件。

采集边界的静态修订建议

以下只示范安全请求边界,属于本稿静态修订,不是原文代码,未执行。调用之前仍须验证 owner/repository 路径,只构造 HTTPS、默认端口、无用户凭据的固定 GitHub API URL。下例仅示范局部边界,不是完整的 URL 安全验证器:

from urllib.parse import urlparse

def github_json(session, url, token):
    if urlparse(url).hostname != "api.github.com":
        raise ValueError("only the GitHub API host is allowed")
    response = session.get(
        url, headers={"Authorization": f"Bearer {token}",
                      "Accept": "application/vnd.github+json"},
        timeout=(3.05, 20), allow_redirects=False)
    response.raise_for_status()
    return response.json()

token = os.environ.get("GITHUB_ACCESS_TOKEN")
if not token:
    raise RuntimeError("GITHUB_ACCESS_TOKEN is required")
# OPENAI_API_KEY 也从受控的运行时 secret 注入读取。

真实实现还须检查 Git tree 的 truncated、Contents 的编码字段/大文件返回形态、限速与响应体大小。仅传必要文件给模型;README、代码注释与文件名都是不可信数据,可能含提示注入文本,不能把它们当任务指令。源入口把结果写到 llms.txt,应拒绝静默覆盖:先写临时文件、核对根目录和 diff,人工确认后再替换。输出只能当待审草稿,不是权威 API 文档。

完整原程序与输出结构

以下按官方教程顺序保留完整签名、RepositoryAnalyzer、三个GitHub采集函数、模型配置/主入口/文件写出以及预期文档结构。原程序只分析元信息,没有执行输出代码或验证API。原始HTTP请求无统一超时/大小/重定向限制,固定main且不检查树截断,环境变量赋占位值会覆盖已有设置;下列代码不能直接视为安全运行版本。占位值并非凭据,没有执行任何请求、模型调用或文件覆盖。

签名定义

import dspy
from typing import List

class AnalyzeRepository(dspy.Signature):
    """Analyze a repository structure and identify key components."""
    repo_url: str = dspy.InputField(desc="GitHub repository URL")
    file_tree: str = dspy.InputField(desc="Repository file structure")
    readme_content: str = dspy.InputField(desc="README.md content")

    project_purpose: str = dspy.OutputField(desc="Main purpose and goals of the project")
    key_concepts: list[str] = dspy.OutputField(desc="List of important concepts and terminology")
    architecture_overview: str = dspy.OutputField(desc="High-level architecture description")

class AnalyzeCodeStructure(dspy.Signature):
    """Analyze code structure to identify important directories and files."""
    file_tree: str = dspy.InputField(desc="Repository file structure")
    package_files: str = dspy.InputField(desc="Key package and configuration files")

    important_directories: list[str] = dspy.OutputField(desc="Key directories and their purposes")
    entry_points: list[str] = dspy.OutputField(desc="Main entry points and important files")
    development_info: str = dspy.OutputField(desc="Development setup and workflow information")

class GenerateLLMsTxt(dspy.Signature):
    """Generate a comprehensive llms.txt file from analyzed repository information."""
    project_purpose: str = dspy.InputField()
    key_concepts: list[str] = dspy.InputField()
    architecture_overview: str = dspy.InputField()
    important_directories: list[str] = dspy.InputField()
    entry_points: list[str] = dspy.InputField()
    development_info: str = dspy.InputField()
    usage_examples: str = dspy.InputField(desc="Common usage patterns and examples")

    llms_txt_content: str = dspy.OutputField(desc="Complete llms.txt file content following the standard format")

RepositoryAnalyzer模块

class RepositoryAnalyzer(dspy.Module):
    def __init__(self):
        super().__init__()
        self.analyze_repo = dspy.ChainOfThought(AnalyzeRepository)
        self.analyze_structure = dspy.ChainOfThought(AnalyzeCodeStructure)
        self.generate_examples = dspy.ChainOfThought("repo_info -> usage_examples")
        self.generate_llms_txt = dspy.ChainOfThought(GenerateLLMsTxt)

    def forward(self, repo_url, file_tree, readme_content, package_files):
        # Analyze repository purpose and concepts
        repo_analysis = self.analyze_repo(
            repo_url=repo_url,
            file_tree=file_tree,
            readme_content=readme_content
        )

        # Analyze code structure
        structure_analysis = self.analyze_structure(
            file_tree=file_tree,
            package_files=package_files
        )

        # Generate usage examples
        usage_examples = self.generate_examples(
            repo_info=f"Purpose: {repo_analysis.project_purpose}\nConcepts: {repo_analysis.key_concepts}"
        )

        # Generate final llms.txt
        llms_txt = self.generate_llms_txt(
            project_purpose=repo_analysis.project_purpose,
            key_concepts=repo_analysis.key_concepts,
            architecture_overview=repo_analysis.architecture_overview,
            important_directories=structure_analysis.important_directories,
            entry_points=structure_analysis.entry_points,
            development_info=structure_analysis.development_info,
            usage_examples=usage_examples.usage_examples
        )

        return dspy.Prediction(
            llms_txt_content=llms_txt.llms_txt_content,
            analysis=repo_analysis,
            structure=structure_analysis
        )

GitHub采集函数

import requests
import os
from pathlib import Path

os.environ["GITHUB_ACCESS_TOKEN"] = "<your_access_token>"

def get_github_file_tree(repo_url):
    """Get repository file structure from GitHub API."""
    # Extract owner/repo from URL
    parts = repo_url.rstrip('/').split('/')
    owner, repo = parts[-2], parts[-1]

    api_url = f"https://api.github.com/repos/{owner}/{repo}/git/trees/main?recursive=1"
    response = requests.get(api_url, headers={
        "Authorization": f"Bearer {os.environ.get('GITHUB_ACCESS_TOKEN')}"
    })

    if response.status_code == 200:
        tree_data = response.json()
        file_paths = [item['path'] for item in tree_data['tree'] if item['type'] == 'blob']
        return '\n'.join(sorted(file_paths))
    else:
        raise Exception(f"Failed to fetch repository tree: {response.status_code}")

def get_github_file_content(repo_url, file_path):
    """Get specific file content from GitHub."""
    parts = repo_url.rstrip('/').split('/')
    owner, repo = parts[-2], parts[-1]

    api_url = f"https://api.github.com/repos/{owner}/{repo}/contents/{file_path}"
    response = requests.get(api_url, headers={
        "Authorization": f"Bearer {os.environ.get('GITHUB_ACCESS_TOKEN')}"
    })

    if response.status_code == 200:
        import base64
        content = base64.b64decode(response.json()['content']).decode('utf-8')
        return content
    else:
        return f"Could not fetch {file_path}"

def gather_repository_info(repo_url):
    """Gather all necessary repository information."""
    file_tree = get_github_file_tree(repo_url)
    readme_content = get_github_file_content(repo_url, "README.md")

    # Get key package files
    package_files = []
    for file_path in ["pyproject.toml", "setup.py", "requirements.txt", "package.json"]:
        try:
            content = get_github_file_content(repo_url, file_path)
            if "Could not fetch" not in content:
                package_files.append(f"=== {file_path} ===\n{content}")
        except:
            continue

    package_files_content = "\n\n".join(package_files)

    return file_tree, readme_content, package_files_content

模型配置与写出入口

def generate_llms_txt_for_dspy():
    # Configure DSPy (use your preferred LM)
    lm = dspy.LM(model="gpt-4o-mini")
    dspy.configure(lm=lm)
    os.environ["OPENAI_API_KEY"] = "<YOUR OPENAI KEY>"

    # Initialize our analyzer
    analyzer = RepositoryAnalyzer()

    # Gather DSPy repository information
    repo_url = "https://github.com/stanfordnlp/dspy"
    file_tree, readme_content, package_files = gather_repository_info(repo_url)

    # Generate llms.txt
    result = analyzer(
        repo_url=repo_url,
        file_tree=file_tree,
        readme_content=readme_content,
        package_files=package_files
    )

    return result

# Run the generation
if __name__ == "__main__":
    result = generate_llms_txt_for_dspy()

    # Save the generated llms.txt
    with open("llms.txt", "w") as f:
        f.write(result.llms_txt_content)

    print("Generated llms.txt file!")
    print("\nPreview:")
    print(result.llms_txt_content[:500] + "...")

原预期llms.txt结构

# DSPy: Programming Language Models

## Project Overview
DSPy is a framework for programming—rather than prompting—language models...

## Key Concepts
- **Modules**: Building blocks for LM programs
- **Signatures**: Input/output specifications  
- **Teleprompters**: Optimization algorithms
- **Predictors**: Core reasoning components

## Architecture
- `/dspy/`: Main package directory
  - `/adapters/`: Input/output format handlers
  - `/clients/`: LM client interfaces
  - `/predict/`: Core prediction modules
  - `/teleprompt/`: Optimization algorithms

## Usage Examples
1. **Building a Classifier**: Using DSPy, a user can define a modular classifier that takes in text data and categorizes it into predefined classes. The user can specify the classification logic declaratively, allowing for easy adjustments and optimizations.
2. **Creating a RAG Pipeline**: A developer can implement a retrieval-augmented generation pipeline that first retrieves relevant documents based on a query and then generates a coherent response using those documents. DSPy facilitates the integration of retrieval and generation components seamlessly.
3. **Optimizing Prompts**: Users can leverage DSPy to create a system that automatically optimizes prompts for language models based on performance metrics, improving the quality of responses over time without manual intervention.
4. **Implementing Agent Loops**: A user can design an agent loop that continuously interacts with users, learns from feedback, and refines its responses, showcasing the self-improving capabilities of the DSPy framework.
5. **Compositional Code**: Developers can write compositional code that allows different modules of the AI system to interact with each other, enabling complex workflows that can be easily modified and extended.

原预期输出以DSPy项目目的、Modules/Signatures/Teleprompters/Predictors概念、adapters/clients/predict/teleprompt目录为骨架,并举分类器、RAG、提示优化、agent循环和组合代码五类用途。这只是示例结构,不证明生成内容正确。原文的后续方向包括多仓库分析、其他文档格式、文档质量指标、交互式Web界面,当前程序没有实现这些扩展。

来源、修改与许可

来源:Generating llms.txt,DSPy官方教程,©2026 DSPy;中文翻译、解释和静态安全注释:未完纪。原片段在此完整保留,前文明确标出的请求边界示例属于编辑修改。以下完整保留DSPy项目的MIT版权与许可通知;用户仓库内容、第三方文档及生成输出须分别核查,不因DSPy使用MIT而自动获得其他材料权利。

MIT License 全文

MIT License

Copyright (c) 2023 Stanford Future Data Systems

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容