Playbook 中的错误处理

当 Ansible 收到命令的非零返回码,或模块报告失败时,默认会停止在该主机上执行任务,并继续在其他主机上执行。不过,在某些情况下,你可能希望采用不同的行为。有时非零返回码表示成功;有时你希望一台主机上的失败让所有主机都停止执行。Ansible 提供了一些工具和设置,用于处理这些情况,帮助你获得所需的行为、输出和报告。

忽略失败的命令

默认情况下,某台主机上的任务失败后,Ansible 会停止在该主机上执行后续任务。你可以使用 ignore_errors,让执行在失败后继续。

- name: Do not count this as a failure
  ansible.builtin.command: /bin/false
  ignore_errors: true

ignore_errors 指令仅在任务可以运行,并且返回“failed”值时生效。它不会使 Ansible 忽略未定义变量错误、连接失败、执行问题(例如缺少软件包)或语法错误。

忽略主机不可达错误

2.7 版本新增。

你可以使用 ignore_unreachable 关键字,忽略因主机实例处于“UNREACHABLE”状态而产生的任务失败。Ansible 会忽略该任务的错误,但仍继续尝试在这台不可达主机上执行后续任务。例如,在任务级别:

- name: This executes, fails, and the failure is ignored
  ansible.builtin.command: /bin/true
  ignore_unreachable: true

- name: This executes, fails, and ends the play for this host
  ansible.builtin.command: /bin/true

在 Playbook 级别:

- hosts: all
  ignore_unreachable: true
  tasks:
  - name: This executes, fails, and the failure is ignored
    ansible.builtin.command: /bin/true

  - name: This executes, fails, and ends the play for this host
    ansible.builtin.command: /bin/true
    ignore_unreachable: false

重置不可达主机

如果 Ansible 无法连接到某台主机,它会将该主机标记为“UNREACHABLE”,并从本次运行的活跃主机列表中移除。你可以使用 meta: clear_host_errors 重新激活所有主机,让后续任务能够再次尝试连接它们。

处理器与失败

Ansible 在每个 play 结束时运行处理器(handler)。如果某个任务通知了一个处理器,但同一个 play 中稍后的任务失败了,默认情况下,该处理器不会在这台主机上运行;这可能使主机处于意料之外的状态。例如,一个任务更新配置文件,并通知处理器重启某个服务;如果同一个 play 中后续任务失败,配置文件可能已经更改,但服务却不会重启。

你可以通过命令行选项 --force-handlers、在 play 中加入 force_handlers: True,或在 ansible.cfg 中加入 force_handlers = True,改变这一行为。强制运行处理器时,Ansible 会在所有主机上运行所有已收到通知的处理器,包括任务曾失败的主机。(注意,某些错误仍可能阻止处理器运行,例如主机变得不可达。)

定义失败

Ansible 允许你使用 failed_when 条件,为每个任务定义“失败”的含义。与 Ansible 中的所有条件一样,由多个 failed_when 条件组成的列表会通过隐含的 and 连接,即只有所有条件都成立,任务才会失败。如果希望任何一个条件成立时都触发失败,就必须将条件写成一个字符串,并显式使用 or 运算符。

例如,当两个条件中的任意一个为真时,让任务失败:

- name: Fail task when either condition is met
  ansible.builtin.command: /usr/bin/example-command
  register: command_result
  failed_when: command_result.rc != 0 or 'ERROR' in command_result.stdout

你可以通过在命令输出中查找某个单词或短语来判断失败:

- name: Fail task when the command error output prints FAILED
  ansible.builtin.command: /usr/bin/example-command -x -y -z
  register: command_result
  failed_when: "'FAILED' in command_result.stderr"

也可以根据返回码判断:

- name: Fail task when both files are identical
  ansible.builtin.raw: diff foo/file1 bar/file2
  register: diff_cmd
  failed_when: diff_cmd.rc == 0 or diff_cmd.rc >= 2

你还可以组合多个失败条件。下列任务会在两个条件都为真时失败:

- name: Check if a file exists in temp and fail task if it does
  ansible.builtin.command: ls /tmp/this_should_not_be_here
  register: result
  failed_when:
    - result.rc == 0
    - '"No such" not in result.stderr'

如果希望只要满足一个条件就让任务失败,可以将 failed_when 定义改为:

failed_when: result.rc == 0 or "No such" not in result.stderr

如果条件太多,无法清晰地放在同一行,可以使用 > 将它们拆成一个多行 YAML 值。

- name: example of many failed_when conditions with OR
  ansible.builtin.shell: "./myBinary"
  register: ret
  failed_when: >
    ("No such file or directory" in ret.stdout) or
    (ret.stderr != '') or
    (ret.rc == 10)

你还可以通过隐含变量 _task 的 result 属性访问任务结果,而无须注册变量。

- name: Fail task when either condition is met
  ansible.builtin.command: /usr/bin/example-command
  failed_when: _task.result.rc != 0 or 'ERROR' in _task.result.stdout

定义“changed”

Ansible 允许你使用 changed_when 条件,定义某个任务何时“更改”了远程节点。你可以根据返回码或输出,决定是否在 Ansible 统计中报告更改,以及是否触发处理器。与 Ansible 中的所有条件一样,由多个 changed_when 条件组成的列表会通过隐含的 and 连接,即只有所有条件都成立,任务才报告更改。如果希望任意一个条件成立就报告更改,必须把条件写成一个字符串,并显式使用 or 运算符。例如:

tasks:

  - name: Report 'changed' when the return code is not equal to 2
    ansible.builtin.shell: /usr/bin/billybass --mode="take me to the river"
    register: bass_result
    changed_when: "bass_result.rc != 2"

  - name: This will never report 'changed' status
    ansible.builtin.shell: wall 'beep'
    changed_when: False

  - name: This task will always report 'changed' status
    ansible.builtin.command: /path/to/command
    changed_when: True

你也可以组合多个条件,覆盖“changed”结果。

- name: Combine multiple conditions to override 'changed' result
  ansible.builtin.command: /bin/fake_command
  register: result
  ignore_errors: True
  changed_when:
    - '"ERROR" in result.stderr'
    - result.rc == 2

与 failed_when 一样,你可以使用隐含变量 _task,避免注册变量:

- name: Combine multiple conditions to override 'changed' result
  ansible.builtin.command: /bin/fake_command
  changed_when:
    - some_msg in _task.result.stdout
    - some_warning not in _task.result.stderr

你可以在条件中引用简单变量,以免重复某些内容,如下例所示:

- name: Example playbook
  hosts: myHosts
  vars:
    log_path: /home/ansible/logfolder/
    log_file: log.log

  tasks:
    - name: Create empty log file
      ansible.builtin.shell: mkdir {{ log_path }} || touch {{ log_path }}{{ log_file }}
      register: tmp
      changed_when:
        - tmp.rc == 0
        - 'tmp.stderr != "mkdir: cannot create directory ‘" ~ log_path ~ "’: File exists"'

更多条件语法示例,参见定义失败。

确保 command 和 shell 成功

command 和 shell 模块会关注返回码。因此,如果某个命令成功时返回的退出码不是零,你可以这样处理:

tasks:
  - name: Run this command and ignore the result
    ansible.builtin.shell: /usr/bin/somecommand || /bin/true

在所有主机上中止 play

有时,你希望单台主机上的失败,或一定比例主机上的失败,让所有主机上的整个 play 中止。可以使用 any_errors_fatal,在第一次失败发生后停止 play 执行。若需要更细粒度的控制,可以使用 max_fail_percentage,在失败主机达到指定比例后中止运行。

第一次错误时中止:any_errors_fatal

如果设置了 any_errors_fatal,并且某个任务返回错误,Ansible 会先在当前批次的所有主机上完成这个致命任务,然后停止在所有主机上执行 play。后续任务和 play 都不会执行。你可以在块中添加 rescue 部分,从致命错误中恢复。any_errors_fatal 可以设置在 play 或块级别。

- hosts: somehosts
  any_errors_fatal: true
  roles:
    - myrole

- hosts: somehosts
  tasks:
    - block:
        - include_tasks: mytasks.yml
      any_errors_fatal: true

当所有任务必须 100% 成功才能继续执行 Playbook 时,可以使用这一功能。例如,你在多个数据中心的机器上运行服务,使用负载均衡器将用户流量转发给服务;在停止服务进行维护前,你希望先禁用所有负载均衡器。为了保证禁用负载均衡器的任务一旦失败,就停止其他所有任务,可以这样写:

---
- hosts: load_balancers_dc_a
  any_errors_fatal: true

  tasks:
    - name: Shut down datacenter 'A'
      ansible.builtin.command: /usr/bin/disable-dc

- hosts: frontends_dc_a

  tasks:
    - name: Stop service
      ansible.builtin.command: /usr/bin/stop-software

    - name: Update software
      ansible.builtin.command: /usr/bin/upgrade-software

- hosts: load_balancers_dc_a

  tasks:
    - name: Start datacenter 'A'
      ansible.builtin.command: /usr/bin/enable-dc

在这个示例中,只有所有负载均衡器都成功禁用,Ansible 才会开始在前端服务器上升级软件。

设置最大失败百分比

默认情况下,只要仍有主机尚未失败,Ansible 就继续执行任务。在滚动更新等场景中,你可能希望失败达到某个阈值时中止 play。为此,可以在 play 上设置最大失败百分比:

---
- hosts: webservers
  max_fail_percentage: 30
  serial: 10

将 max_fail_percentage 与 serial 配合使用时,该设置会作用于每个批次。在上面的示例中,如果第一批(或任意一批)10 台服务器中有超过 3 台失败,就会中止该 play 的剩余部分。

在块中控制错误

你也可以使用块来定义对任务错误的响应。这种做法类似于许多编程语言中的异常处理。详情和示例参见使用块处理错误。

另请参阅

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容