用 PySpark 表函数展开行并按分区汇总

标量函数一次调用返回一个值,用户自定义表函数则可以返回零行、一行或多行。Spark 从 3.5 开始引入 Python UDTF:它可以出现在 SQL 的 FROM 子句中,接受零个或多个标量参数,也可以接受表示整张输入表的表参数。它适合拆分文本、展开日期区间、逐行转换,以及在明确的分区边界内维护状态并输出结果。

本文完整整理 Apache Spark 官方《Python User-defined Table Functions (UDTFs)》。2026-10-05 读取的 latest 页面标识为 PySpark 4.2.0;“从 3.5 引入”不代表本文所有后续功能在 3.5 都可用。尤其是 Arrow 默认行为、动态分析和表参数能力,应以实际部署版本文档为准。所有代码只做静态审核,未启动 Spark。

PySpark UDTF生命周期示意:analyze确定schema和分区,executor实例逐行eval,成功后terminate产生汇总,最后cleanup;各分区状态独立
原创 UDTF 生命周期与分区示意图;不是执行计划或运行截图。

一、先建立会话,再在 executor 上运行表函数

官方示例默认已有 spark。在标准 PySpark 环境中,可先准备会话和常用导入:

from pyspark.sql import SparkSession, Row
from pyspark.sql.functions import lit, udtf

spark = SparkSession.builder.appName("udtf-tutorial").getOrCreate()

这段是本文补充的示例前提。运行环境仍需要满足所用 PySpark 版本的 Python、Java 和依赖要求;集群地址与资源配置按实际环境设置。UDTF 的实例在 executor 一侧创建,不能在类中创建或引用驱动端 SparkSession,否则可能因序列化和执行上下文限制而失败。

二、UDTF 的五个方法及其生命周期

一个 Python UDTF 用类来实现,其中 eval() 是必需方法;__init__()、analyze()、terminate() 与 cleanup() 按需实现。

方法 何时使用 关键约束
__init__ executor 实例创建时初始化状态 实例状态保留到当前输入分区处理结束;不能引用 SparkSession
analyze 分析特定调用的参数,决定输出 schema 和表参数要求 静态方法;接收 AnalyzeArgument,而非运行时普通值
eval 逐个输入行求值 必须实现;可产出零到多行,列数和类型应符合 schema
terminate 正常处理完输入后收尾并可继续产出行 普通异常中止时不会调用;不适合承担唯一的资源清理职责
cleanup 正常或异常结束后的清理 用于释放本实例持有的外部资源;不是输出结果的方法

初始化与状态范围

__init__() 中写入的字段可供后续 eval() 和 terminate() 使用。一个实例会一直存在到它所消费的当前分区输入结束;不同分区的实例状态互相隔离。不要用类字段累加后,就误以为得到了整个 DataFrame 的全局总数。

如果实现了 analyze(),构造方法还可以写成 __init__(self, analyze_result: AnalyzeResult),接收分析结果。可以继承 AnalyzeResult 增加自定义元数据,让分析时只需执行一次的计算传给以后创建的实例,减少重复初始化成本。

eval 的参数与输出

每个标量表达式映射成一个 Python 值;表参数映射成一个 pyspark.sql.Row,字段顺序与输入表对应,名称由查询分析器确定。每消费一个输入行,调用一次 eval();它可以使用明确的参数列表、*args 或 **kwargs。

# 以下是方法写法示例,放在具有匹配输出schema的UDTF类内。
def eval(self, x: int):
    yield (x,)

def eval(self, x: int, y: int):
    yield (x + y, x - y)
    yield (y + x, y - x)

def eval(self, *args):
    if len(args) != 2:
        raise ValueError("需要两个参数")
    x, y = args
    yield (x + y, x - y)
    yield (y + x, y - x)

def eval(self, **kwargs):
    if set(kwargs) != {"x", "y"}:
        raise ValueError("需要x和y参数")
    x, y = kwargs["x"], kwargs["y"]
    yield (x + y, x - y)
    yield (y + x, y - x)

这四个片段是互为替代的写法,不是让同一个类重复定义四次 eval。原文用 assert 验证参数个数,本文改为显式异常,以免 Python 优化模式关闭断言后丢失检查。类型注解帮助阅读,但不是运行时完整的输入校验。

eval() 可以抛出 SkipRestOfInputTableException,表示不再消费当前分区剩余的输入,直接进入 terminate()。其他异常会中止该 UDTF,跳过 terminate(),转入 cleanup(),并把异常传播到查询处理器,使调用查询失败。

收尾输出与最终清理

terminate() 可以输出最后的汇总行,也可以在消费完整个分区后统一构造结果。后者需要尤其注意内存占用:不要为了收尾方便,把无限规模的输入全部保存在 Python 列表中。

# 收尾输出示意:要求输出schema恰好有两列,类型与值匹配。
def terminate(self):
    yield "done", None

# 资源清理示意:conn必须由本实例在受控逻辑中创建。
def cleanup(self):
    conn = getattr(self, "conn", None)
    if conn is not None:
        conn.close()

原文清理片段是直接调用 self.conn.close(),本文增加未初始化时的检查。分布式任务可能失败或重试,因此外部写入、网络请求等副作用还需另外设计幂等性;UDTF 生命周期钩子本身不提供业务级“只执行一次”的承诺。

三、定义输出 schema 和产出行

返回类型决定输出表的字段名、顺序和类型。可以在 udtf 注册时提供静态 schema,也可以由 analyze() 返回动态 schema。静态类型可写为 StructType 或 DDL 字符串:

from pyspark.sql.types import StructType, StringType

schema = StructType().add("c1", StringType())
# 等价的DDL表达:
schema_ddl = "c1 string"

eval() 和 terminate() 通过 yield 产出元组、列表或 Row。一个元素对应 schema 的一列,可以产出多次。三列的几种写法如下:

from pyspark.sql import Row

def eval(self, x, y, z):
    yield (x, y, z)
    yield x, y, z
    yield Row(x, y, z)

def terminate(self):
    yield [self.x, self.y, self.z]

上面前三条语句演示等价的行表达形式;如果全部保留,会实际产出三行。只有一列时,元组末尾的逗号不能省:

def eval(self, x):
    yield (x,)

原文 Row 示例的 def eval(self, x, y, z) 缺少冒号,本文已补齐,并使用公开常见的 from pyspark.sql import Row 导入路径。

四、注册到 SQL:拆词与 LATERAL 调用

先定义一个接受字符串并逐词输出的表函数:

from pyspark.sql.functions import udtf

@udtf(returnType="word: string")
class WordSplitter:
    def eval(self, text: str):
        for word in text.split(" "):
            yield (word.strip(),)

spark.udtf.register("split_words", WordSplitter)
spark.sql("SELECT * FROM split_words('hello world')").show()

这个调用输出一列 word,两行分别是 hello、world。代码保留原文的 split(" ") 语义:连续空格可能产生空字符串,它不等同于不带参数的 split();空值也没有在这个简短示例中处理。实际业务可根据需求明确过滤空词或接收 NULL 的策略。

当函数输入来自前面的 FROM 项时,使用 LATERAL:

spark.sql(
    "SELECT * FROM VALUES ('Hello World'), ('Apache Spark') t(text), "
    "LATERAL split_words(text)"
).show()

LATERAL 允许引用前面表项的列与别名。结果有 text、word 两列:Hello World 对应 Hello 与 World 两行;Apache Spark 对应 Apache 与 Spark 两行。输入行因表函数的多行输出而展开,不能按标量 UDF 的“一行进、一值出”来理解。

五、Arrow 优化及版本差异

Apache Arrow 是内存中的列式数据格式,Spark 用它提高 Java 与 Python 进程之间的数据传输效率。此次读取的 4.2 文档明确写明:从 Spark 4.2 起,Python UDTF 默认启用 Arrow。若要禁用,可设置:

spark.conf.set("spark.sql.execution.pythonUDTF.arrow.enabled", "false")

若需要显式启用,可把同一配置设为 "true",或者在声明时传入 useArrow=True:

from pyspark.sql.functions import udtf

@udtf(returnType="c1: int, c2: int", useArrow=True)
class PlusOne:
    def eval(self, x: int):
        yield x, x + 1

原文指出,一个输入行生成较大结果表时,Arrow 可能带来收益。“可能”不等于对任意数据或所有 UDTF 都保证更快。Arrow 依赖、类型转换和版本兼容性应参考Apache Arrow in PySpark。不要把 4.2 的默认值套用到旧版本。

六、标量参数示例:平方数、日期展开和计数

两种方式创建平方数表函数

可以先定义普通类,再调用 udtf() 包装:

from pyspark.sql.functions import lit, udtf

class SquareNumbers:
    def eval(self, start: int, end: int):
        for num in range(start, end + 1):
            yield (num, num * num)

square_num = udtf(SquareNumbers, returnType="num: int, squared: int")
square_num(lit(1), lit(3)).show()

也可以直接在类上使用装饰器,作为前一写法的替代:

@udtf(returnType="num: int, squared: int")
class SquareNumbers:
    def eval(self, start: int, end: int):
        for num in range(start, end + 1):
            yield (num, num * num)

SquareNumbers(lit(1), lit(3)).show()

两种调用的结果都是 (1, 1)、(2, 4)、(3, 9),范围包含终点。示例声明的是 Spark int;扩大输入范围时,平方结果可能超出该类型范围,不能因为 Python 整数可增长就忽略返回 schema。生产代码还需对区间长度、空值和反向区间规定清楚。

将日期区间展开成每日一行

from datetime import datetime, timedelta
from pyspark.sql.functions import lit, udtf

@udtf(returnType="date: string")
class DateExpander:
    def eval(self, start_date: str, end_date: str):
        current = datetime.strptime(start_date, "%Y-%m-%d")
        end = datetime.strptime(end_date, "%Y-%m-%d")
        while current <= end:
            yield (current.strftime("%Y-%m-%d"),)
            current += timedelta(days=1)

DateExpander(lit("2023-02-25"), lit("2023-03-01")).show()

结果依次是 2023-02-25、02-26、02-27、02-28 和 03-01,包含起止两日。列名虽叫 date,声明类型实际是字符串,不是 Spark DateType。无效日期会导致解析异常,跨度过大则会显著放大行数。例子用于展示展开逻辑,接收外部输入时应增加可接受的最大跨度和空值策略。

用 terminate 输出分区计数

from pyspark.sql.functions import udtf

@udtf(returnType="cnt: int")
class CountUDTF:
    def __init__(self):
        self.count = 0

    def eval(self, x: int):
        self.count += 1

    def terminate(self):
        yield self.count,

spark.udtf.register("count_udtf", CountUDTF)
spark.sql(
    "SELECT * FROM range(0, 10, 1, 1), LATERAL count_udtf(id)"
).show()
spark.sql(
    "SELECT * FROM range(0, 10, 1, 2), LATERAL count_udtf(id)"
).show()

第一条 range 只有一个分区,原文结果为 id=9, cnt=10。第二条有两个分区,原文结果为 (id=4, cnt=5) 和 (id=9, cnt=5)。这是收尾输出与当前处理分区的关联效果,不是每个 id 都获得全局计数。若输入很大,还应考虑把 cnt 改成适合规模的 bigint。

七、把整张表作为参数传入

UDTF 可以在标量参数之外接受一个表参数;官方当前指南规定,同一调用最多只有一个这种表参数。SQL 形式可以是 TABLE(t),也可以是 TABLE(SELECT a, b, c FROM t) 或包含连接的子查询。UDTF 逐行接收一个 Row,并非一次收到整个 DataFrame。

from pyspark.sql import Row
from pyspark.sql.functions import udtf

@udtf(returnType="id: int")
class FilterUDTF:
    def eval(self, row: Row):
        if row["id"] > 5:
            yield row["id"],

spark.udtf.register("filter_udtf", FilterUDTF)
spark.sql(
    "SELECT * FROM filter_udtf(TABLE(SELECT * FROM range(10)))"
).show()

该例输出 id 为 6、7、8、9 的四行。输入没有通过条件时,eval() 不 yield 即可输出零行。使用列名索引前,应确保输入 schema 包含所需字段,明确处理 NULL 的语义。

八、明确输入分区与顺序,再做有状态汇总

在表参数之后加 PARTITION BY,可以保证同一分区表达式组合的输入行交给同一个 UDTF 实例。表达式不限于列名,也可使用字符串长度、日期月份或组合值。也可以用 WITH SINGLE PARTITION,要求所有输入都由一个实例消费。

需要有序处理时,在 PARTITION BY 或 WITH SINGLE PARTITION 后添加 ORDER BY。这是分区内部送入 eval 的顺序要求,与最外层 SQL 的结果排序不同。

按 a 分区并求 b 的最大值

原文建立持久表时先执行 DROP TABLE IF EXISTS,最后再删除。为避免覆盖已有同名表,下面改用专用临时视图;示例输入与三种调用语义保持一致。最大值的初始状态也从原文的 0 改成 None,否则全为负数时可能错误输出 0。

from pyspark.sql import Row
from pyspark.sql.functions import udtf

@udtf(returnType="a: string, b: int")
class PartitionMax:
    def __init__(self):
        self.key = None
        self.maximum = None

    def eval(self, row: Row):
        self.key = row["a"]
        value = row["b"]
        if value is not None:
            if self.maximum is None or value > self.maximum:
                self.maximum = value

    def terminate(self):
        if self.key is not None:
            yield self.key, self.maximum

spark.udtf.register("partition_max_wwj2673", PartitionMax)
spark.createDataFrame(
    [("abc", 2), ("abc", 4), ("def", 6), ("def", 8)],
    "a string, b int"
).createOrReplaceTempView("wwj2673_values")

spark.sql("""
    SELECT *
    FROM partition_max_wwj2673(
        TABLE(wwj2673_values) PARTITION BY a ORDER BY b
    )
    ORDER BY 1
""").show()

输入四行是 (abc,2)、(abc,4)、(def,6)、(def,8)。按 a 分区后,两个实例分别得到 abc 和 def 的行,输出 (abc,4)、(def,8)。本例输入的 a 均非 NULL;若业务允许 NULL 分区键,应另设“是否收到行”的状态,不能沿用当前 key is not None 判断。

按表达式分区

spark.sql("""
    SELECT *
    FROM partition_max_wwj2673(
        TABLE(wwj2673_values) PARTITION BY LENGTH(a) ORDER BY b
    )
    ORDER BY 1
""").show()

abc 和 def 的长度都为 3,因此四行进入同一个实例。按 b 升序处理后,最后保存的 key 是 def,最大值是 8,结果为 (def,8)。这里的 a 列表示最后消费行的键,并不表示“所有长度为 3 的记录都属于 def”。更一般的分组汇总应输出真正的分区表达式,或按业务定义保存所需键。

全部输入放入单一分区

spark.sql("""
    SELECT *
    FROM partition_max_wwj2673(
        TABLE(wwj2673_values) WITH SINGLE PARTITION ORDER BY b
    )
    ORDER BY 1
""").show()

# 只清理本示例创建的临时视图。
spark.catalog.dropTempView("wwj2673_values")

这个调用同样输出 (def,8)。单分区让全体输入交给同一个实例,但也限制了这一阶段的并行能力;不能为了获得全局状态就不计输入规模地使用。若排序键存在并列值,又依赖“最后一行”的含义,还需要稳定的次级排序或明确的并列规则。

九、动态 analyze:根据调用参数决定 schema

静态返回 schema 适合大多数示例。如果输出列取决于调用时的参数,可以不在 udtf 上提供 returnType,改由静态 analyze() 返回 AnalyzeResult。这里的“分析”发生在查询规划阶段:接收的是参数的类型和元信息,不能假定可以读取任意运行时行值。

AnalyzeArgument 提供什么

字段 含义
dataType 传入表达式的数据类型;表参数是描述输入列的 StructType
value 可在分析阶段取得的常量值;表参数和非常量表达式通常为 None
isTable 是否为表参数
isConstantExpression 是否为字面量或可常量折叠的标量表达式

因此,不能直接对一个 AnalyzeArgument 调用字符串的 split();应先验证类型和是否为可取得值的常量,再读取 .value。NULL 常量的 value 也可能为 None,不能只用 value 是否为 None 代替所有类型检查。

AnalyzeResult 提供什么

字段 请求的行为
schema 输出表的 StructType,定义列名、顺序和类型
withSinglePartition 默认 False;为 True 时要求输入表由一个实例消费
partitionBy PartitioningColumn 序列;按表达式值重分区
orderBy OrderingColumn 序列;指定各分区内的消费顺序
select SelectedColumn 序列;让 Catalyst 先计算表达式,只把选中的字段按指定顺序交给 UDTF

分区、排序和选择要求是查询规划器的输入,不是让 Python 函数自己在内存中完成一次全表重分区。它们可能引入 shuffle、排序和数据移动,性能仍取决于规模及分区分布。

示例:一个常量单词对应一列

原文先给出按参数名、*args、**kwargs 的三种分析方法,但其中直接把参数对象当字符串使用,还出现 kwargs 分支误写为 args["text"]。下面保留同一构想并修正,先集中定义分析逻辑:

from pyspark.sql.functions import udtf
from pyspark.sql.types import StructType, StringType
from pyspark.sql.udtf import AnalyzeArgument, AnalyzeResult

def analyze_word_columns(arg: AnalyzeArgument) -> AnalyzeResult:
    if arg.isTable or arg.dataType != StringType():
        raise TypeError("需要字符串标量参数")
    if not arg.isConstantExpression or not isinstance(arg.value, str):
        raise ValueError("输出列数需要可在分析阶段确定的非NULL字符串常量")
    words = arg.value.split(" ")
    if len(words) > 1024:
        raise ValueError("本示例最多生成1024列")
    schema = StructType()
    for index, _ in enumerate(words):
        schema = schema.add(f"word_{index}", StringType())
    return AnalyzeResult(schema=schema)

@udtf
class WideWords:
    @staticmethod
    def analyze(text: AnalyzeArgument) -> AnalyzeResult:
        return analyze_word_columns(text)

    def eval(self, text: str):
        yield tuple(text.split(" "))

spark.udtf.register("wide_words_wwj2673", WideWords)
spark.sql("SELECT * FROM wide_words_wwj2673('hello world')").show()

这个例子按静态推导会有 word_0 和 word_1 两列,一行内容为 hello、world。1024 列只是本文为演示添加的规模约束,不是 Spark 的通用限制。与前面的 WordSplitter 不同,这里改变的是列数,不是行数。

如果需要位置参数或关键字参数形式,可把类中的静态方法替换成下面任一版本,eval 仍需与函数调用的参数约定匹配:

# 替代写法一:位置参数。
@staticmethod
def analyze(*args) -> AnalyzeResult:
    if len(args) != 1:
        raise ValueError("只接受一个参数")
    return analyze_word_columns(args[0])

# 替代写法二:命名参数。
@staticmethod
def analyze(**kwargs) -> AnalyzeResult:
    if set(kwargs) != {"text"}:
        raise ValueError("只接受名为text的参数")
    return analyze_word_columns(kwargs["text"])

这些是替代方法片段,不能在同一个类里重复定义后指望同时生效。是否使用关键字调用及其具体 SQL 形式,应按目标 Spark 版本支持情况核对。

把分析结果中的自定义元数据传给实例

原文还展示固定输出 schema、额外记录单词数和英文冠词数的方案。它的示意代码包含不匹配引号,并对生成器调用 len()。下面修正后展示元数据类与分析方法;字段设置默认值,避免与继承的 dataclass 默认字段顺序发生冲突。

from dataclasses import dataclass
from pyspark.sql.types import StructType, StringType, IntegerType
from pyspark.sql.udtf import AnalyzeArgument, AnalyzeResult

@dataclass
class WordMetadata(AnalyzeResult):
    num_words: int = 0
    num_articles: int = 0

def analyze_metadata(text: AnalyzeArgument) -> WordMetadata:
    # 重用前面的参数与常量检查,结果列在这里采用固定schema。
    analyze_word_columns(text)
    words = text.value.split(" ")
    return WordMetadata(
        schema=(StructType()
                .add("word", StringType())
                .add("total", IntegerType())),
        num_words=len(words),
        num_articles=sum(1 for word in words if word in {"a", "an", "the"}),
    )

# 在UDTF类内可采用下面的构造方法接收结果:
def __init__(self, analyze_result: WordMetadata):
    self.num_words = analyze_result.num_words
    self.num_articles = analyze_result.num_articles

实际类中的 analyze 应返回 analyze_metadata(...) 的结果;eval 则按 word、total 两列的含义产生数据。原文的目的在于示范元数据传递,本身没有给出完整业务输出逻辑,所以这里也不把这个片段标成可独立执行的完整 UDTF。冠词统计保留大小写敏感、以空格分词的原语义。

让 analyze 请求选列、按月份分区并按日期排序

原文最后一个动态分析例子希望根据输入表的 date 列月份分组,按 date 排序,并只传递 date 与 word 长度。源码中既有引号缺失,也把月份输出类型写成 DateType,却用月份表达式进行分组。下面给出修正后的分析方法片段,明确月份为整数、最长词为长度:

from pyspark.sql.types import StructType, IntegerType, DateType, StringType
from pyspark.sql.udtf import (
    AnalyzeArgument, AnalyzeResult,
    PartitioningColumn, OrderingColumn, SelectedColumn,
)

# 放在动态schema UDTF类中;eval和terminate要与返回两列对应。
@staticmethod
def analyze(table: AnalyzeArgument) -> AnalyzeResult:
    if not table.isTable:
        raise TypeError("需要一个表参数")
    fields = table.dataType
    if "date" not in fields.names or "word" not in fields.names:
        raise ValueError("输入表需要date和word列")
    if fields["date"].dataType != DateType():
        raise TypeError("date需要是DateType")
    if fields["word"].dataType != StringType():
        raise TypeError("word需要是StringType")
    return AnalyzeResult(
        schema=(StructType()
                .add("month", IntegerType())
                .add("longest_word", IntegerType())),
        partitionBy=[PartitioningColumn("extract(month from date)")],
        orderBy=[OrderingColumn("date")],
        select=[
            SelectedColumn("date"),
            SelectedColumn(name="length(word)", alias="length_word"),
        ],
    )

该片段保留原文按“月份数字”分区的语义,因此不同年份的同一个月份会被放在一起。如果业务要求按年月汇总,应修改分区表达式及输出 schema;不能在不说明的情况下把这两种需求混用。选列之后,eval 收到的字段是 date 与 length_word,而不是原始 word。示例仍需要相应的逐行状态与收尾实现,且应定义日期、词字段为 NULL 时的行为。

十、SQL 子句与 analyze 字段的对应关系

前面在 SQL 中写出的输入要求,也可以由 UDTF 自己在分析阶段提出。例如,调用方只写 SELECT * FROM f(TABLE(t)),函数仍可通过返回下列字段请求相应计划:

调用方显式写法 analyze 中的对应字段
TABLE(t) PARTITION BY a partitionBy=[PartitioningColumn("a")]
TABLE(t) WITH SINGLE PARTITION withSinglePartition=True
分区要求后的 ORDER BY b orderBy=[OrderingColumn("b")]
TABLE(SELECT a FROM t) select=[SelectedColumn("a")]

这张对应表说明能力关系,不表示互相冲突的分区要求可以随意同时指定。原文叙述里用小写 true 表示布尔值,Python 代码应写为 True。当 UDTF 对输入顺序或分区有硬性依赖时,把要求集中放进 analyze 有助于避免调用方遗漏,但仍需让调用者理解它对 shuffle、并行度和输出语义的影响。

十一、使用前的静态审核结论

本篇保留了官方完整教程中的生命周期、schema、SQL 注册、Arrow、所有标量与表参数示例、三种分区调用和动态分析章节。主线优先使用静态 schema,动态 analyze 作为完整进阶章节保留;原文存在错误的片段没有被直接包装成已验证代码。

主要修正包括:SparkSession 前提;AnalyzeArgument 与常量 value 的区别;静态 analyze 不带实例 self;位置与关键字参数笔误;缺失或不匹配引号;生成器计数;Row 示例缺少冒号;继承 dataclass 的默认字段;月份类型与选列别名;最大值初始化;以及持久 DROP/CREATE 表替换为专用临时视图。对于替代方法和不完整业务片段,正文明确标出其使用位置与边界。

仍需在目标集群验证 schema 转换、Arrow 依赖、分区规划、空值、错误传播、数据倾斜和大规模输入。Python UDTF 可能显著放大输出行数;有状态实现也可能积累过多内存。优先确认是否已有合适的内建 SQL 函数,再决定是否使用自定义表函数。未发现静态问题不等于没有运行时缺陷。

来源与署名

原文:Apache Spark 官方《Python User-defined Table Functions (UDTFs)》;维护与发布方为 Apache Software Foundation,页面未见个人署名。版本化代码示例中的 ASF 许可声明应随再利用保留;原页面明确标注:Copyright 2026 The Apache Software Foundation, Licensed under the Apache License, Version 2.0。中文翻译及编辑标注:未完纪,2026-10-05;完整许可证与适用通知保留于文末。

本文配图为原创技术示意,代码均未执行,结果表述来自原文示例或明确的静态推导,不是本次 Spark 运行记录。

版权与许可全文

以下保留本页涉及的来源材料或示例代码的版权、许可条件与免责声明;各自适用范围依原声明。中文翻译及编辑标注:未完纪,2026-10-05。

Apache-2.0

                                 Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/

   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION

   1. Definitions.

      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.

      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.

      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.

      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.

      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.

      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.

      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).

      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.

      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."

      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by Licensor and
      subsequently incorporated within the Work.

   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.

   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.

   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and

      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and

      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding those notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.

      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.

   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.

   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.

   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.

   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.

   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.

   END OF TERMS AND CONDITIONS

   APPENDIX: How to apply the Apache License to your work.

      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "[]"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. We also recommend that a
      file or class name and description of purpose be included on the
      same "printed page" as the copyright notice for easier
      identification within third-party archives.

   Copyright [yyyy] [name of copyright owner]

   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

       http://www.apache.org/licenses/LICENSE-2.0

   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
   limitations under the License.

Spark-NOTICE

Apache Spark
Copyright 2014 and onwards The Apache Software Foundation.

This product includes software developed at
The Apache Software Foundation (http://www.apache.org/).


Export Control Notice
---------------------

This distribution includes cryptographic software. The country in which you currently reside may have
restrictions on the import, possession, use, and/or re-export to another country, of encryption software.
BEFORE using any encryption software, please check your country's laws, regulations and policies concerning
the import, possession, or use, and re-export of encryption software, to see if this is permitted. See
<http://www.wassenaar.org/> for more information.

The U.S. Government Department of Commerce, Bureau of Industry and Security (BIS), has classified this
software as Export Commodity Control Number (ECCN) 5D002.C.1, which includes information security software
using or performing cryptographic functions with asymmetric algorithms. The form and manner of this Apache
Software Foundation distribution makes it eligible for export under the License Exception ENC Technology
Software Unrestricted (TSU) exception (see the BIS Export Administration Regulations, Section 740.13) for
both object code and source code.

The following provides more details on the included cryptographic software:

This software uses Apache Commons Crypto (https://commons.apache.org/proper/commons-crypto/) to
support authentication, and encryption and decryption of data sent across the network between
services.


Metrics
Copyright 2010-2013 Coda Hale and Yammer, Inc.

This product includes software developed by Coda Hale and Yammer, Inc.

This product includes code derived from the JSR-166 project (ThreadLocalRandom, Striped64,
LongAdder), which was released with the following comments:

    Written by Doug Lea with assistance from members of JCP JSR-166
    Expert Group and released to the public domain, as explained at
    http://creativecommons.org/publicdomain/zero/1.0/
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容