把文本特征处理与分类训练串成 Spark Pipeline

译编来源:Apache Spark:ML Pipelines。维护方:Apache Software Foundation,原页未注明个人作者。本次读取的页面标为 Spark 4.2.0;Python 示例固定到官方仓库 v4.2.0。

一个文本分类任务通常不只包含分类器:先把文本拆成词,再把词转换成数值特征,最后根据特征和标签训练模型。ML Pipelines 为这些步骤提供基于 DataFrame 的统一高层接口,使处理流程能一起配置、拟合和复用。

本文完整整理原页概念与行为边界,代码统一采用 Python 版本;原页还提供功能相近的 Scala 和 Java 写法,其完整示例仍可从原文进入。本例只有四条人工训练记录,说明的是接口如何衔接,不能据此判断文本分类质量。

Spark流水线训练时依次分词、哈希特征、拟合逻辑回归;推断时复用同样特征步骤并由拟合模型输出概率和预测。
训练与推断使用同一组特征处理阶段。原创示意:未完纪编辑部,依据 Apache Spark ML Pipelines。

DataFrame、Transformer 与 Estimator

DataFrame 是这里的数据集载体:一张表中可以同时保留文本、数值特征向量、真实标签以及预测结果。Spark SQL 的基本和结构化类型之外,DataFrame 还支持机器学习用的 Vector 类型。DataFrame 可由 RDD 显式或隐式建立,列使用名称定位,例如 text、features 和 label。

Transformer 将一个 DataFrame 变成另一个 DataFrame,通常增加一列或多列。特征转换器读文本列、追加向量列;训练好的模型读特征列、追加预测列。两者都通过 transform() 工作。

Estimator 接收数据并拟合,fit() 返回 Model,而 Model 也是 Transformer。例如 LogisticRegression.fit() 训练后得到 LogisticRegressionModel。原文将 transform() 和 fit() 描述为无状态操作;每个组件实例都有唯一 ID,可用来精确关联其参数。

流水线怎样训练,又怎样推断

Pipeline 按顺序保存多个 PipelineStage,每个阶段可以是 Transformer 或 Estimator。训练时,Transformer 对当前数据调用 transform();Estimator 对当前数据调用 fit(),得到对应的 Transformer。如果后面还有需要训练的阶段,再用这个已拟合 Transformer 处理数据并传下去。

在文本例子中,Tokenizer 把 text 拆成 words;HashingTF 从 words 生成 features;LogisticRegression 使用特征和标签训练分类器。Pipeline 自身是 Estimator,pipeline.fit(training) 的结果是 PipelineModel。

PipelineModel 的阶段数与原 Pipeline 相同,但其中原来的 Estimator 已替换为训练好的 Transformer。对测试数据调用 model.transform(test),就会按顺序进行相同特征处理和模型预测。这样能避免训练和推断分别编写预处理时产生的不一致。

流水线不局限于直线结构。阶段的输入列和输出列隐式定义数据依赖,只要组成有向无环图,就可安排非线性流程;此时阶段数组必须按拓扑顺序排列。Pipeline 和 PipelineModel 根据 DataFrame schema 在运行前检查输入类型,不能把这种检查理解成编译期类型保证。

每个阶段还应是独立实例。不要把同一个 myHashingTF 对象在同一流水线放两次;若确实需要两个 HashingTF 阶段,应分别创建实例,使 ID 不同。

一个完整的 Python 文本分类例子

下面保留官方 pipeline_example.py 的完整代码和许可头,未改算法或参数。它建立 SparkSession、创建内嵌训练数据、拟合流水线,再为不含标签的测试文本生成结果,最后停止会话。需要已安装并配置匹配版本的 Spark/PySpark 和运行环境;本文未安装依赖、启动 Spark 或运行训练。

#
# Licensed to the Apache Software Foundation (ASF) under one or more
# contributor license agreements.  See the NOTICE file distributed with
# this work for additional information regarding copyright ownership.
# The ASF licenses this file to You under the Apache License, Version 2.0
# (the "License"); you may not use this file except in compliance with
# the License.  You may obtain a copy of the License at
#
#    http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#

"""
Pipeline Example.
"""

# $example on$
from pyspark.ml import Pipeline
from pyspark.ml.classification import LogisticRegression
from pyspark.ml.feature import HashingTF, Tokenizer
# $example off$
from pyspark.sql import SparkSession

if __name__ == "__main__":
    spark = SparkSession\
        .builder\
        .appName("PipelineExample")\
        .getOrCreate()

    # $example on$
    # Prepare training documents from a list of (id, text, label) tuples.
    training = spark.createDataFrame([
        (0, "a b c d e spark", 1.0),
        (1, "b d", 0.0),
        (2, "spark f g h", 1.0),
        (3, "hadoop mapreduce", 0.0)
    ], ["id", "text", "label"])

    # Configure an ML pipeline, which consists of three stages: tokenizer, hashingTF, and lr.
    tokenizer = Tokenizer(inputCol="text", outputCol="words")
    hashingTF = HashingTF(inputCol=tokenizer.getOutputCol(), outputCol="features")
    lr = LogisticRegression(maxIter=10, regParam=0.001)
    pipeline = Pipeline(stages=[tokenizer, hashingTF, lr])

    # Fit the pipeline to training documents.
    model = pipeline.fit(training)

    # Prepare test documents, which are unlabeled (id, text) tuples.
    test = spark.createDataFrame([
        (4, "spark i j k"),
        (5, "l m n"),
        (6, "spark hadoop spark"),
        (7, "apache hadoop")
    ], ["id", "text"])

    # Make predictions on test documents and print columns of interest.
    prediction = model.transform(test)
    selected = prediction.select("id", "text", "probability", "prediction")
    for row in selected.collect():
        rid, text, prob, prediction = row
        print(
            "(%d, %s) --> prob=%s, prediction=%f" % (
                rid, text, str(prob), prediction
            )
        )
    # $example off$

    spark.stop()

id 用于识别记录,text 是输入文本,label 只出现在训练数据中。probability 是模型输出的类别概率向量,prediction 是预测类别。示例打印这些值,但本文没有生成或编造实际概率输出。

这里的 Python 示例没有显式设置 HashingTF 的 numFeatures,沿用当前版本默认值;原页 Scala 和 Java 的流水线示例则显式设为 1000。这一差异应在跨语言迁移时主动统一。训练和推断必须保持分词、特征列和哈希维度一致。

统一参数 API 与覆盖关系

所有 Transformer 和 Estimator 使用统一的参数 API。一个 Param 有名称和说明,ParamMap 则保存参数与值的对应关系。可以直接为实例设参数,例如构造逻辑回归时传 maxIter=10、regParam=0.01,或调用 setter;也可在 fit()、transform() 时传入 ParamMap。后者会覆盖实例此前设置的同名参数。

参数属于具体实例,而不是仅靠字符串名称定位。两个逻辑回归对象可以分别设置自己的 maxIter。Python 的参数映射使用字典,原文例子先设迭代次数 20,随后覆盖成 30,再指定正则参数 0.1、阈值 0.55,并把概率输出列改名为 myProbability。

# 摘自官方参数例;依赖原例已创建的 lr 和 training。
paramMap = {lr.maxIter: 20}
paramMap[lr.maxIter] = 30
paramMap.update({lr.regParam: 0.1, lr.threshold: 0.55})
paramMap2 = {lr.probabilityCol: "myProbability"}
paramMapCombined = paramMap.copy()
paramMapCombined.update(paramMap2)
model2 = lr.fit(training, paramMapCombined)

因此检查训练参数时,应以最终传入的映射为准。原例还用 explainParams() 查看参数说明和默认值,用 extractParamMap() 查看模型使用的参数,并对新的向量数据调用 transform()。输出选择的是改名后的 myProbability,继续读取 probability 就会与当前设置不符。完整参数示例见下方附录及官方版本化文件。

保存模型时同时考虑版本

Spark 1.6 为 Pipeline API 加入了模型导入导出;原文说明,自 Spark 2.3 起,基于 DataFrame 的 spark.ml 和 pyspark.ml 已完整覆盖持久化。Scala、Java、Python 的持久化可跨语言使用;R 在原文描述中仍采用调整过的格式,R 保存的模型只能由 R 加载,相关限制指向 SPARK-15572。

持久化可保存训练好的 PipelineModel,也可保存尚未训练的 Pipeline。原页 Scala 示例使用 write.overwrite().save(...) 写到两个固定的 /tmp 路径,然后用 PipelineModel.load(...) 加载。编辑风险提示:overwrite() 明确允许覆盖既有目标;迁移到业务代码时应使用独立且已确认的输出目录,不能照抄固定路径去覆盖正在使用的模型。

原文对兼容性的承诺有边界:跨大版本尽力兼容,但不保证;小版本和补丁版本支持向后加载兼容。持久化文件的内部格式本身不承诺稳定。模型行为在小版本和补丁版本之间通常保持一致,但错误修复可能改变行为。破坏兼容性的变动应在发行说明中列出;这些说法都不能代替你对目标版本的验证。

流水线的另一个作用是把超参数优化放进统一流程。自动模型选择的完整用法在原文链接的 ML Tuning Guide 中;本文不把四条记录扩展成未做过的交叉验证或准确率报告。

静态审核与使用边界

本次只读取文档、源码与许可,未执行示例。内嵌字符串作为数据送入 DataFrame,没有在这段代码中发现动态执行、命令拼接或硬编码凭据;这不表示整个运行环境或 Spark 没有安全问题。

collect() 把所选数据全部取回驱动进程;示例只有四条测试数据,换成大数据集可能耗尽驱动内存。结果还会打印原始文本,因此不应把敏感业务文本直接送入可共享日志。SparkSession.getOrCreate() 使用环境中的会话或配置,运行前需要确认连接目标是隔离测试环境,不能把文中调用理解成必然只在本机运行。

本稿的编辑提示不改变原示例。没有加载任何外部模型、访问生产集群,也没有测试性能、精度、跨语言模型加载或版本兼容性。

原文与代码归属:Apache Spark / The Apache Software Foundation。Apache Spark:Copyright 2014 and onwards The Apache Software Foundation。本示例代码依 Apache License 2.0 提供,完整 LICENSE 和上游 NOTICE 见下方。中文译编、图示和标注的审核说明:未完纪编辑部。

完整参数示例:Estimator、Transformer 与 Param

#
# Licensed to the Apache Software Foundation (ASF) under one or more
# contributor license agreements.  See the NOTICE file distributed with
# this work for additional information regarding copyright ownership.
# The ASF licenses this file to You under the Apache License, Version 2.0
# (the "License"); you may not use this file except in compliance with
# the License.  You may obtain a copy of the License at
#
#    http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
#

"""
Estimator Transformer Param Example.
"""
# $example on$
from pyspark.ml.linalg import Vectors
from pyspark.ml.classification import LogisticRegression
# $example off$
from pyspark.sql import SparkSession

if __name__ == "__main__":
    spark = SparkSession\
        .builder\
        .appName("EstimatorTransformerParamExample")\
        .getOrCreate()

    # $example on$
    # Prepare training data from a list of (label, features) tuples.
    training = spark.createDataFrame([
        (1.0, Vectors.dense([0.0, 1.1, 0.1])),
        (0.0, Vectors.dense([2.0, 1.0, -1.0])),
        (0.0, Vectors.dense([2.0, 1.3, 1.0])),
        (1.0, Vectors.dense([0.0, 1.2, -0.5]))], ["label", "features"])

    # Create a LogisticRegression instance. This instance is an Estimator.
    lr = LogisticRegression(maxIter=10, regParam=0.01)
    # Print out the parameters, documentation, and any default values.
    print("LogisticRegression parameters:\n" + lr.explainParams() + "\n")

    # Learn a LogisticRegression model. This uses the parameters stored in lr.
    model1 = lr.fit(training)

    # Since model1 is a Model (i.e., a transformer produced by an Estimator),
    # we can view the parameters it used during fit().
    # This prints the parameter (name: value) pairs, where names are unique IDs for this
    # LogisticRegression instance.
    print("Model 1 was fit using parameters: ")
    print(model1.extractParamMap())

    # We may alternatively specify parameters using a Python dictionary as a paramMap
    paramMap = {lr.maxIter: 20}
    paramMap[lr.maxIter] = 30  # Specify 1 Param, overwriting the original maxIter.
    # Specify multiple Params.
    paramMap.update({lr.regParam: 0.1, lr.threshold: 0.55})  # type: ignore

    # You can combine paramMaps, which are python dictionaries.
    # Change output column name
    paramMap2 = {lr.probabilityCol: "myProbability"}
    paramMapCombined = paramMap.copy()
    paramMapCombined.update(paramMap2)  # type: ignore

    # Now learn a new model using the paramMapCombined parameters.
    # paramMapCombined overrides all parameters set earlier via lr.set* methods.
    model2 = lr.fit(training, paramMapCombined)
    print("Model 2 was fit using parameters: ")
    print(model2.extractParamMap())

    # Prepare test data
    test = spark.createDataFrame([
        (1.0, Vectors.dense([-1.0, 1.5, 1.3])),
        (0.0, Vectors.dense([3.0, 2.0, -0.1])),
        (1.0, Vectors.dense([0.0, 2.2, -1.5]))], ["label", "features"])

    # Make predictions on test data using the Transformer.transform() method.
    # LogisticRegression.transform will only use the 'features' column.
    # Note that model2.transform() outputs a "myProbability" column instead of the usual
    # 'probability' column since we renamed the lr.probabilityCol parameter previously.
    prediction = model2.transform(test)
    result = prediction.select("features", "label", "myProbability", "prediction") \
        .collect()

    for row in result:
        print("features=%s, label=%s -> prob=%s, prediction=%s"
              % (row.features, row.label, row.myProbability, row.prediction))
    # $example off$

    spark.stop()

Apache License 2.0 全文

                                 Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/

   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION

   1. Definitions.

      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.

      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.

      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.

      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.

      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.

      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.

      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).

      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.

      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."

      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by Licensor and
      subsequently incorporated within the Work.

   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.

   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.

   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and

      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and

      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding those notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.

      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.

   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.

   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.

   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.

   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.

   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.

   END OF TERMS AND CONDITIONS

   APPENDIX: How to apply the Apache License to your work.

      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "[]"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. We also recommend that a
      file or class name and description of purpose be included on the
      same "printed page" as the copyright notice for easier
      identification within third-party archives.

   Copyright [yyyy] [name of copyright owner]

   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

       http://www.apache.org/licenses/LICENSE-2.0

   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
   limitations under the License.


------------------------------------------------------------------------------------
This product bundles various third-party components under other open source licenses.
This section summarizes those components and their licenses. See licenses/
for text of these licenses.


Apache Software Foundation License 2.0
--------------------------------------

common/network-common/src/main/java/org/apache/spark/network/util/LimitedInputStream.java
core/src/main/java/org/apache/spark/util/collection/TimSort.java
core/src/main/resources/org/apache/spark/ui/static/bootstrap*
core/src/main/resources/org/apache/spark/ui/static/vis*
connector/spark-ganglia-lgpl/src/main/java/com/codahale/metrics/ganglia/GangliaReporter.java
core/src/main/resources/org/apache/spark/ui/static/d3-flamegraph.min.js
core/src/main/resources/org/apache/spark/ui/static/d3-flamegraph.css
mllib-local/src/main/scala/scala/collection/compat/package.scala

Python Software Foundation License
----------------------------------

python/pyspark/loose_version.py

BSD 3-Clause
------------

python/lib/py4j-*-src.zip
python/pyspark/cloudpickle/*.py
python/pyspark/join.py

The CSS style for the navigation sidebar of the documentation was originally
submitted by Óscar Nájera for the scikit-learn project. The scikit-learn project
is distributed under the 3-Clause BSD license.


MIT License
-----------

core/src/main/resources/org/apache/spark/ui/static/dagre-d3.min.js
core/src/main/resources/org/apache/spark/ui/static/*dataTables*
core/src/main/resources/org/apache/spark/ui/static/graphlib-dot.min.js
core/src/main/resources/org/apache/spark/ui/static/jquery*
core/src/main/resources/org/apache/spark/ui/static/sorttable.js
docs/js/vendor/bootstrap*
docs/js/vendor/anchor.min.js
docs/js/vendor/jquery*
docs/js/vendor/modernizer*
docs/js/vendor/docsearch.min.js

ISC License
-----------

core/src/main/resources/org/apache/spark/ui/static/d3.min.js


Creative Commons CC0 1.0 Universal Public Domain Dedication
-----------------------------------------------------------
(see LICENSE-CC0.txt)

data/mllib/images/kittens/29.5.a_b_EGDP022204.jpg
data/mllib/images/kittens/54893.jpg
data/mllib/images/kittens/DP153539.jpg
data/mllib/images/kittens/DP802813.jpg
data/mllib/images/multi-channel/chr30.4.184.jpg

上游 NOTICE

Apache Spark
Copyright 2014 and onwards The Apache Software Foundation.

This product includes software developed at
The Apache Software Foundation (http://www.apache.org/).


Export Control Notice
---------------------

This distribution includes cryptographic software. The country in which you currently reside may have
restrictions on the import, possession, use, and/or re-export to another country, of encryption software.
BEFORE using any encryption software, please check your country's laws, regulations and policies concerning
the import, possession, or use, and re-export of encryption software, to see if this is permitted. See
<http://www.wassenaar.org/> for more information.

The U.S. Government Department of Commerce, Bureau of Industry and Security (BIS), has classified this
software as Export Commodity Control Number (ECCN) 5D002.C.1, which includes information security software
using or performing cryptographic functions with asymmetric algorithms. The form and manner of this Apache
Software Foundation distribution makes it eligible for export under the License Exception ENC Technology
Software Unrestricted (TSU) exception (see the BIS Export Administration Regulations, Section 740.13) for
both object code and source code.

The following provides more details on the included cryptographic software:

This software uses Apache Commons Crypto (https://commons.apache.org/proper/commons-crypto/) to
support authentication, and encryption and decryption of data sent across the network between
services.


Metrics
Copyright 2010-2013 Coda Hale and Yammer, Inc.

This product includes software developed by Coda Hale and Yammer, Inc.

This product includes code derived from the JSR-166 project (ThreadLocalRandom, Striped64,
LongAdder), which was released with the following comments:

    Written by Doug Lea with assistance from members of JCP JSR-166
    Expert Group and released to the public domain, as explained at
    http://creativecommons.org/publicdomain/zero/1.0/
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容