当 Ansible 收到命令的非零返回码,或模块报告失败时,默认会停止在该主机上执行任务,并继续在其他主机上执行。不过,在某些情况下,你可能希望采用不同的行为。有时非零返回码表示成功;有时你希望一台主机上的失败让所有主机都停止执行。Ansible 提供了一些工具和设置,用于处理这些情况,帮助你获得所需的行为、输出和报告。
忽略失败的命令
默认情况下,某台主机上的任务失败后,Ansible 会停止在该主机上执行后续任务。你可以使用 ignore_errors,让执行在失败后继续。
- name: Do not count this as a failure
ansible.builtin.command: /bin/false
ignore_errors: true
ignore_errors 指令仅在任务可以运行,并且返回“failed”值时生效。它不会使 Ansible 忽略未定义变量错误、连接失败、执行问题(例如缺少软件包)或语法错误。
忽略主机不可达错误
2.7 版本新增。
你可以使用 ignore_unreachable 关键字,忽略因主机实例处于“UNREACHABLE”状态而产生的任务失败。Ansible 会忽略该任务的错误,但仍继续尝试在这台不可达主机上执行后续任务。例如,在任务级别:
- name: This executes, fails, and the failure is ignored
ansible.builtin.command: /bin/true
ignore_unreachable: true
- name: This executes, fails, and ends the play for this host
ansible.builtin.command: /bin/true
在 Playbook 级别:
- hosts: all
ignore_unreachable: true
tasks:
- name: This executes, fails, and the failure is ignored
ansible.builtin.command: /bin/true
- name: This executes, fails, and ends the play for this host
ansible.builtin.command: /bin/true
ignore_unreachable: false
重置不可达主机
如果 Ansible 无法连接到某台主机,它会将该主机标记为“UNREACHABLE”,并从本次运行的活跃主机列表中移除。你可以使用 meta: clear_host_errors 重新激活所有主机,让后续任务能够再次尝试连接它们。
处理器与失败
Ansible 在每个 play 结束时运行处理器(handler)。如果某个任务通知了一个处理器,但同一个 play 中稍后的任务失败了,默认情况下,该处理器不会在这台主机上运行;这可能使主机处于意料之外的状态。例如,一个任务更新配置文件,并通知处理器重启某个服务;如果同一个 play 中后续任务失败,配置文件可能已经更改,但服务却不会重启。
你可以通过命令行选项 --force-handlers、在 play 中加入 force_handlers: True,或在 ansible.cfg 中加入 force_handlers = True,改变这一行为。强制运行处理器时,Ansible 会在所有主机上运行所有已收到通知的处理器,包括任务曾失败的主机。(注意,某些错误仍可能阻止处理器运行,例如主机变得不可达。)
定义失败
Ansible 允许你使用 failed_when 条件,为每个任务定义“失败”的含义。与 Ansible 中的所有条件一样,由多个 failed_when 条件组成的列表会通过隐含的 and 连接,即只有所有条件都成立,任务才会失败。如果希望任何一个条件成立时都触发失败,就必须将条件写成一个字符串,并显式使用 or 运算符。
例如,当两个条件中的任意一个为真时,让任务失败:
- name: Fail task when either condition is met
ansible.builtin.command: /usr/bin/example-command
register: command_result
failed_when: command_result.rc != 0 or 'ERROR' in command_result.stdout
你可以通过在命令输出中查找某个单词或短语来判断失败:
- name: Fail task when the command error output prints FAILED
ansible.builtin.command: /usr/bin/example-command -x -y -z
register: command_result
failed_when: "'FAILED' in command_result.stderr"
也可以根据返回码判断:
- name: Fail task when both files are identical
ansible.builtin.raw: diff foo/file1 bar/file2
register: diff_cmd
failed_when: diff_cmd.rc == 0 or diff_cmd.rc >= 2
你还可以组合多个失败条件。下列任务会在两个条件都为真时失败:
- name: Check if a file exists in temp and fail task if it does
ansible.builtin.command: ls /tmp/this_should_not_be_here
register: result
failed_when:
- result.rc == 0
- '"No such" not in result.stderr'
如果希望只要满足一个条件就让任务失败,可以将 failed_when 定义改为:
failed_when: result.rc == 0 or "No such" not in result.stderr
如果条件太多,无法清晰地放在同一行,可以使用 > 将它们拆成一个多行 YAML 值。
- name: example of many failed_when conditions with OR
ansible.builtin.shell: "./myBinary"
register: ret
failed_when: >
("No such file or directory" in ret.stdout) or
(ret.stderr != '') or
(ret.rc == 10)
你还可以通过隐含变量 _task 的 result 属性访问任务结果,而无须注册变量。
- name: Fail task when either condition is met
ansible.builtin.command: /usr/bin/example-command
failed_when: _task.result.rc != 0 or 'ERROR' in _task.result.stdout
定义“changed”
Ansible 允许你使用 changed_when 条件,定义某个任务何时“更改”了远程节点。你可以根据返回码或输出,决定是否在 Ansible 统计中报告更改,以及是否触发处理器。与 Ansible 中的所有条件一样,由多个 changed_when 条件组成的列表会通过隐含的 and 连接,即只有所有条件都成立,任务才报告更改。如果希望任意一个条件成立就报告更改,必须把条件写成一个字符串,并显式使用 or 运算符。例如:
tasks:
- name: Report 'changed' when the return code is not equal to 2
ansible.builtin.shell: /usr/bin/billybass --mode="take me to the river"
register: bass_result
changed_when: "bass_result.rc != 2"
- name: This will never report 'changed' status
ansible.builtin.shell: wall 'beep'
changed_when: False
- name: This task will always report 'changed' status
ansible.builtin.command: /path/to/command
changed_when: True
你也可以组合多个条件,覆盖“changed”结果。
- name: Combine multiple conditions to override 'changed' result
ansible.builtin.command: /bin/fake_command
register: result
ignore_errors: True
changed_when:
- '"ERROR" in result.stderr'
- result.rc == 2
与 failed_when 一样,你可以使用隐含变量 _task,避免注册变量:
- name: Combine multiple conditions to override 'changed' result
ansible.builtin.command: /bin/fake_command
changed_when:
- some_msg in _task.result.stdout
- some_warning not in _task.result.stderr
你可以在条件中引用简单变量,以免重复某些内容,如下例所示:
- name: Example playbook
hosts: myHosts
vars:
log_path: /home/ansible/logfolder/
log_file: log.log
tasks:
- name: Create empty log file
ansible.builtin.shell: mkdir {{ log_path }} || touch {{ log_path }}{{ log_file }}
register: tmp
changed_when:
- tmp.rc == 0
- 'tmp.stderr != "mkdir: cannot create directory ‘" ~ log_path ~ "’: File exists"'
更多条件语法示例,参见定义失败。
确保 command 和 shell 成功
command 和 shell 模块会关注返回码。因此,如果某个命令成功时返回的退出码不是零,你可以这样处理:
tasks:
- name: Run this command and ignore the result
ansible.builtin.shell: /usr/bin/somecommand || /bin/true
在所有主机上中止 play
有时,你希望单台主机上的失败,或一定比例主机上的失败,让所有主机上的整个 play 中止。可以使用 any_errors_fatal,在第一次失败发生后停止 play 执行。若需要更细粒度的控制,可以使用 max_fail_percentage,在失败主机达到指定比例后中止运行。
第一次错误时中止:any_errors_fatal
如果设置了 any_errors_fatal,并且某个任务返回错误,Ansible 会先在当前批次的所有主机上完成这个致命任务,然后停止在所有主机上执行 play。后续任务和 play 都不会执行。你可以在块中添加 rescue 部分,从致命错误中恢复。any_errors_fatal 可以设置在 play 或块级别。
- hosts: somehosts
any_errors_fatal: true
roles:
- myrole
- hosts: somehosts
tasks:
- block:
- include_tasks: mytasks.yml
any_errors_fatal: true
当所有任务必须 100% 成功才能继续执行 Playbook 时,可以使用这一功能。例如,你在多个数据中心的机器上运行服务,使用负载均衡器将用户流量转发给服务;在停止服务进行维护前,你希望先禁用所有负载均衡器。为了保证禁用负载均衡器的任务一旦失败,就停止其他所有任务,可以这样写:
---
- hosts: load_balancers_dc_a
any_errors_fatal: true
tasks:
- name: Shut down datacenter 'A'
ansible.builtin.command: /usr/bin/disable-dc
- hosts: frontends_dc_a
tasks:
- name: Stop service
ansible.builtin.command: /usr/bin/stop-software
- name: Update software
ansible.builtin.command: /usr/bin/upgrade-software
- hosts: load_balancers_dc_a
tasks:
- name: Start datacenter 'A'
ansible.builtin.command: /usr/bin/enable-dc
在这个示例中,只有所有负载均衡器都成功禁用,Ansible 才会开始在前端服务器上升级软件。
设置最大失败百分比
默认情况下,只要仍有主机尚未失败,Ansible 就继续执行任务。在滚动更新等场景中,你可能希望失败达到某个阈值时中止 play。为此,可以在 play 上设置最大失败百分比:
---
- hosts: webservers
max_fail_percentage: 30
serial: 10
将 max_fail_percentage 与 serial 配合使用时,该设置会作用于每个批次。在上面的示例中,如果第一批(或任意一批)10 台服务器中有超过 3 台失败,就会中止该 play 的剩余部分。
在块中控制错误
你也可以使用块来定义对任务错误的响应。这种做法类似于许多编程语言中的异常处理。详情和示例参见使用块处理错误。
另请参阅
- Ansible Playbook:Playbook 简介。
- 通用技巧:Playbook 的技巧和窍门。
- 条件语句:Playbook 中的条件语句。
- 使用变量:关于变量的完整介绍。
- 交流:有疑问?需要帮助?想分享想法?请访问 Ansible 交流指南。











暂无评论内容