nbconvert:移除单元格、输入或输出

移除单元格、输入或输出

将 Notebook 转换为其他格式时,可以使用预处理器移除单元格的一部分,也可以移除整个单元格。Notebook 本身保持不变,只有导出的输出会去除指定部分。主要有以下两种实现方式。

使用单元格标签移除单元格的部分内容

控制移除哪些内容的最直接方式是使用单元格标签。标签是保存在每个单元格“tag”字段中的单字符串元数据片段。TagRemovePreprocessor 可用于移除输入、输出或整个单元格。

例如,下面的配置对单元格的各个部分使用不同标签,配合 HTMLExporter 分别移除。在这里演示的是 nbconvert Python API。

from traitlets.config import Config
import nbformat as nbf
from nbconvert.exporters import HTMLExporter
from nbconvert.preprocessors import TagRemovePreprocessor

# Setup config
c = Config()

# Configure tag removal - be sure to tag your cells to remove  using the
# words remove_cell to remove cells. You can also modify the code to use
# a different tag word
c.TagRemovePreprocessor.remove_cell_tags = ("remove_cell",)
c.TagRemovePreprocessor.remove_all_outputs_tags = ("remove_output",)
c.TagRemovePreprocessor.remove_input_tags = ("remove_input",)
c.TagRemovePreprocessor.enabled = True

# Configure and run out exporter
c.HTMLExporter.preprocessors = ["nbconvert.preprocessors.TagRemovePreprocessor"]

exporter = HTMLExporter(config=c)
exporter.register_preprocessor(TagRemovePreprocessor(config=c), True)

# Configure and run our exporter - returns a tuple - first element with html,
# second with notebook metadata
output = HTMLExporter(config=c).from_filename("your-notebook-file-path.ipynb")

# Write to output html file
with open("your-output-file-name.html", "w") as f:
    f.write(output[0])

下面的补充示例演示如何通过命令行接口移除带有特定标签的单元格。

jupyter nbconvert mynotebook.ipynb --TagRemovePreprocessor.enabled=True --TagRemovePreprocessor.remove_cell_tags remove_cell

根据单元格内容使用正则表达式移除单元格

有时候,你希望根据单元格的内容而不是标签进行移除。这时可以使用 RegexRemovePreprocessor。

初始化此预处理器时,传入单个 patterns 配置,它是一个字符串列表。预处理器逐一检查各单元格的内容是否匹配 patterns 中的任一字符串;如果匹配任一模式,就会从导出的 Notebook 中移除该单元格。

例如,执行以下命令,将 Notebook 转换为 HTML,并移除只包含空白字符的单元格:

jupyter nbconvert --RegexRemovePreprocessor.patterns="['\s*\Z']" mynotebook.ipynb

该命令行参数将模式列表设为 '\s*\Z',它匹配任意数量的空白字符以及字符串结尾。

https://regex101.com/ 提供交互式正则表达式指南,请确保选择 Python 方言。Python 官方正则表达式文档请参阅 https://docs.python.org/library/re.html。


原文:Removing cells, inputs, or outputs。来源:Jupyter 官方文档。

© Jupyter Development Team。nbconvert 文档和源码遵循项目 BSD 许可;原始许可见项目仓库。

原始版权与许可

原始许可证

BSD 3-Clause License

- Copyright (c) 2001-2015, IPython Development Team
- Copyright (c) 2015-, Jupyter Development Team

All rights reserved.

Redistribution and use in source and binary forms, with or without
modification, are permitted provided that the following conditions are met:

1. Redistributions of source code must retain the above copyright notice, this
   list of conditions and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice,
   this list of conditions and the following disclaimer in the documentation
   and/or other materials provided with the distribution.

3. Neither the name of the copyright holder nor the names of its
   contributors may be used to endorse or promote products derived from
   this software without specific prior written permission.
THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容