把自定义与第三方类型接入 Pydantic

本文根据 Pydantic Validation 官方文档的 Types 页面整理并中文化。原文页面标题为“Types”,该页面未标注个人作者姓名;本稿标题根据本篇分配元数据拟定。

  • 来源:https://pydantic.dev/docs/validation/latest/concepts/types/
  • 主题:自定义类型、第三方类型适配、CoreSchema 与泛型
  • 版本说明:来源地址指向 latest 文档。正文没有锁定某个 Pydantic 发布版本;文中提到的 Python 版本适用范围按原页面标示保留。
  • 许可说明:已读取的页面正文未列出文章转载许可信息。

Pydantic 怎样理解类型

Pydantic 根据字段类型决定怎样验证输入、怎样序列化结果。Python 内置类型和标准库类型(如 int、str、date)通常可以直接使用;可以控制严格模式,也可以添加约束。Pydantic 和 pydantic-extra-types 还提供额外类型,例如 SecretStr。自定义类型部分说明了这些类型可复用的集成模式。

如果现有类型已经能表达数据结构,优先使用标准类型并添加约束。需要把校验、序列化或 JSON Schema 行为封装成可重复使用的类型时,可以用 Annotated。只有需要更深层控制、或要让没有 Pydantic 集成的外部类参与验证时,才考虑底层 CoreSchema 钩子。

用 Annotated 定义可复用的约束

以正整数为例:将 int 与 Field(gt=0) 组合成一个类型别名,再用 TypeAdapter 独立验证该类型,不必先声明 BaseModel。

from typing import Annotated

from pydantic import Field, TypeAdapter, ValidationError

PositiveInt = Annotated[int, Field(gt=0)]

ta = TypeAdapter(PositiveInt)

print(ta.validate_python(1))
# 1

try:
    ta.validate_python(-1)
except ValidationError as exc:
    print(exc)

传入 1 可通过验证;传入 -1 会产生大于 0 的约束错误。若希望约束定义不依赖 Pydantic,也可以使用 annotated-types:

from typing import Annotated

from annotated_types import Gt

PositiveInt = Annotated[int, Gt(0)]

给类型增加验证、序列化和 JSON Schema

Pydantic 导出的 Annotated 标记可以为既有类型增加或覆盖验证、序列化和 JSON Schema 行为。下面把浮点数四舍五入到一位小数;序列化时格式化为科学计数法字符串,并把序列化模式的 JSON Schema 声明为字符串。

from typing import Annotated

from pydantic import (
    AfterValidator,
    PlainSerializer,
    TypeAdapter,
    WithJsonSchema,
)

TruncatedFloat = Annotated[
    float,
    AfterValidator(lambda x: round(x, 1)),
    PlainSerializer(lambda x: f'{x:.1e}', return_type=str),
    WithJsonSchema({'type': 'string'}, mode='serialization'),
]

ta = TypeAdapter(TruncatedFloat)

value = 1.02345
assert value != 1.0
assert ta.validate_python(value) == 1.0
assert ta.dump_json(value) == b'"1.0e+00"'
assert ta.json_schema(mode='validation') == {'type': 'number'}
assert ta.json_schema(mode='serialization') == {'type': 'string'}

这展示了两条不同路径:验证时仍以数字为类型,应用 AfterValidator;序列化时由 PlainSerializer 输出字符串,JSON Schema 也按序列化模式声明为字符串。

将约束用于泛型类型

Annotated 元数据也可以用于类型变量。以下定义让列表最多包含四项,并让列表元素必须大于 0:

from typing import Annotated, TypeVar

from annotated_types import Gt, Len
from pydantic import TypeAdapter, ValidationError

T = TypeVar('T')

ShortList = Annotated[list[T], Len(max_length=4)]

ta = TypeAdapter(ShortList[int])
assert ta.validate_python([1, 2, 3, 4]) == [1, 2, 3, 4]

try:
    ta.validate_python([1, 2, 3, 4, 5])
except ValidationError as exc:
    print(exc)

PositiveList = list[Annotated[T, Gt(0)]]

ta = TypeAdapter(PositiveList[float])
values = ta.validate_python([1.0])
assert type(values[0]) is float

try:
    ta.validate_python([-1.0])
except ValidationError as exc:
    print(exc)

ShortList 限制列表长度;PositiveList 将元素约束附着在 T 上,因此可以把同一结构用于不同的具体元素类型。

命名类型别名与 JSON Schema 定义

直接把类型赋给变量会产生隐式类型别名。运行时,Pydantic 无法知道变量的别名名称;当同一类型在模型中多次出现时,生成的 JSON Schema 定义可能重复,而且隐式递归别名通常无法正常工作。

Pydantic v2.11 起完整支持命名类型别名。对 Python 3.9 及以上版本,可用 typing_extensions.TypeAliasType;Python 3.12 起也可使用 type 语句。

下面使用 TypeAliasType 命名“正整数列表”,并在两个字段上复用:

from typing import Annotated

from annotated_types import Gt
from typing_extensions import TypeAliasType

from pydantic import BaseModel

PositiveIntList = TypeAliasType(
    'PositiveIntList',
    list[Annotated[int, Gt(0)]],
)

class Model(BaseModel):
    x: PositiveIntList
    y: PositiveIntList

print(Model.model_json_schema())

生成的 Schema 会将 PositiveIntList 放入 $defs,并让 x、y 通过 $ref 指向同一份定义,避免把定义重复写在每个字段下。Python 3.12 的对应语法是:

from typing import Annotated

from annotated_types import Gt
from pydantic import BaseModel

type PositiveIntList = list[Annotated[int, Gt(0)]]

class Model(BaseModel):
    x: PositiveIntList
    y: PositiveIntList

别名内元数据的边界

类型别名适合封装类型自身的约束和 JSON 元数据,但 Pydantic 不会把别名内部的字段级元数据应用到使用别名的字段上。例如,在别名内设置 Field(default=1)、alias 或 deprecated 不受支持。文档说明,若要支持这些字段专属选项,就必须提前展开检查别名的值,也会妨碍 Pydantic 将该别名保留为 JSON Schema 中可复用的定义。

命名泛型别名

命名别名也可以包含类型变量。Python 3.9 及以上版本可用 TypeAliasType 并声明 type_params:

from typing import Annotated, TypeVar

from annotated_types import Len
from typing_extensions import TypeAliasType

T = TypeVar('T')

ShortList = TypeAliasType(
    'ShortList',
    Annotated[list[T], Len(max_length=4)],
    type_params=(T,),
)

Python 3.12 及以上可以写成:

from typing import Annotated

from annotated_types import Len

type ShortList[T] = Annotated[list[T], Len(max_length=4)]

用命名别名定义递归类型

递归类型应使用命名别名。Pydantic 无法支持隐式递归别名的一个原因是,无法可靠地跨模块解析其前向注解。

Python 3.9 及以上可以用 TypeAliasType。这里用字符串表达尚未定义完成的递归引用:

from typing import Union

from typing_extensions import TypeAliasType

from pydantic import TypeAdapter

Json = TypeAliasType(
    'Json',
    'Union[dict[str, Json], list[Json], str, int, float, bool, None]',
)

ta = TypeAdapter(Json)
print(ta.json_schema())

Python 3.12 的新 type 语法会延迟计算别名值,可以直接写递归联合:

from pydantic import TypeAdapter

type Json = dict[str, Json] | list[Json] | str | int | float | bool | None

ta = TypeAdapter(Json)
print(ta.json_schema())

两种定义都会生成含递归引用的 JSON Schema。Pydantic 也提供 JsonValue 类型作为便利选择。

用 CoreSchema 深度定制类型

如果 Annotated、Field、BeforeValidator、AfterValidator 等高层方式已能满足需求,应优先使用这些方式。__get_pydantic_core_schema__ 是更底层的 Pydantic V2 API,用于控制 Pydantic 如何生成 pydantic-core 的验证 Schema;文档提醒,这个 API 较新,未来更可能调整。

此钩子可以定义在自定义类型上,也可以定义在用于 Annotated 的元数据类上。它接收 source_type 和 handler:source_type 不一定等于当前类,泛型尤其如此;handler 可以委托给下一个 Annotated 元数据处理器,或委托给 Pydantic 的内部 Schema 生成器。最简单的实现是调用 handler(source_type) 并原样返回,也可以先改变类型、再调整返回的 CoreSchema,或完全自行构造 Schema。

在自定义类型上实现钩子

下面的 Username 继承 str,并在验证后构造 Username 实例。这与 Pydantic V1 的 __get_validators__ 用途相近:

from typing import Any

from pydantic_core import CoreSchema, core_schema

from pydantic import GetCoreSchemaHandler, TypeAdapter

class Username(str):
    @classmethod
    def __get_pydantic_core_schema__(
        cls, source_type: Any, handler: GetCoreSchemaHandler
    ) -> CoreSchema:
        return core_schema.no_info_after_validator_function(cls, handler(str))

ta = TypeAdapter(Username)
result = ta.validate_python('abc')
assert isinstance(result, Username)
assert result == 'abc'

这里 handler(str) 先生成字符串的验证 Schema;其后处理函数再调用 Username 构造实例。

把验证逻辑作为 Annotated 元数据

有时并不需要定义一个新的运行时类型,只希望在原类型上附加校验。例如,下面用 dataclass 元数据实现与 AfterValidator 类似的后置验证:

from dataclasses import dataclass
from typing import Annotated, Any, Callable

from pydantic_core import CoreSchema, core_schema

from pydantic import BaseModel, GetCoreSchemaHandler

@dataclass(frozen=True)
class MyAfterValidator:
    func: Callable[[Any], Any]

    def __get_pydantic_core_schema__(
        self, source_type: Any, handler: GetCoreSchemaHandler
    ) -> CoreSchema:
        return core_schema.no_info_after_validator_function(
            self.func, handler(source_type)
        )

Username = Annotated[str, MyAfterValidator(str.lower)]

class Model(BaseModel):
    name: Username

assert Model(name='ABC').name == 'abc'

dataclass 标为 frozen=True 后可哈希;否则,像 Username | None 这样的联合类型可能报错。静态类型检查器仍会把 Username 看作 str,而不会视为独立类型。

适配没有 Pydantic 集成的第三方类型

第三方库的类可能既不能修改,也没有为 Pydantic 提供 CoreSchema。可在独立标记类上实现 Schema 钩子,再把标记类作为 Annotated 元数据包在外部类型周围。下面定义的行为是:

  • 输入整数时,创建 ThirdPartyType 实例,并把整数存入 x;
  • 输入已有的 ThirdPartyType 实例时,保留实例及其值;
  • 其他不能转为整数的输入不通过验证;
  • 序列化始终输出整数;
  • JSON Schema 将该字段呈现为整数。
from typing import Annotated, Any

from pydantic_core import core_schema

from pydantic import (
    BaseModel,
    GetCoreSchemaHandler,
    GetJsonSchemaHandler,
    ValidationError,
)
from pydantic.json_schema import JsonSchemaValue

class ThirdPartyType:
    x: int

    def __init__(self):
        self.x = 0

class _ThirdPartyTypePydanticAnnotation:
    @classmethod
    def __get_pydantic_core_schema__(
        cls,
        _source_type: Any,
        _handler: GetCoreSchemaHandler,
    ) -> core_schema.CoreSchema:
        def validate_from_int(value: int) -> ThirdPartyType:
            result = ThirdPartyType()
            result.x = value
            return result

        from_int_schema = core_schema.chain_schema(
            [
                core_schema.int_schema(),
                core_schema.no_info_plain_validator_function(validate_from_int),
            ]
        )

        return core_schema.json_or_python_schema(
            json_schema=from_int_schema,
            python_schema=core_schema.union_schema(
                [
                    core_schema.is_instance_schema(ThirdPartyType),
                    from_int_schema,
                ]
            ),
            serialization=core_schema.plain_serializer_function_ser_schema(
                lambda instance: instance.x
            ),
        )

    @classmethod
    def __get_pydantic_json_schema__(
        cls, _core_schema: core_schema.CoreSchema, handler: GetJsonSchemaHandler
    ) -> JsonSchemaValue:
        return handler(core_schema.int_schema())

PydanticThirdPartyType = Annotated[
    ThirdPartyType, _ThirdPartyTypePydanticAnnotation
]

class Model(BaseModel):
    third_party_type: PydanticThirdPartyType

from_int = Model(third_party_type=1)
assert isinstance(from_int.third_party_type, ThirdPartyType)
assert from_int.third_party_type.x == 1
assert from_int.model_dump() == {'third_party_type': 1}

instance = ThirdPartyType()
instance.x = 10

from_instance = Model(third_party_type=instance)
assert isinstance(from_instance.third_party_type, ThirdPartyType)
assert from_instance.third_party_type.x == 10
assert from_instance.model_dump() == {'third_party_type': 10}

try:
    Model(third_party_type='a')
except ValidationError as exc:
    print(exc)

assert Model.model_json_schema() == {
    'properties': {
        'third_party_type': {'title': 'Third Party Type', 'type': 'integer'}
    },
    'required': ['third_party_type'],
    'title': 'Model',
    'type': 'object',
}

CoreSchema 分别定义了 JSON 输入和 Python 输入的验证路径。两条路径都可以用整数解析 Schema;Python 路径还先检查是否已经是第三方类型实例。序列化器把实例投影成其 x 整数值。JSON Schema 钩子调用 handler,并将整数 Schema 作为该字段的公开表示。

这种适配方式还可用于 Pandas、NumPy 等第三方类型。核心技巧是用轻量的 Annotated 包装,而无需把 Pydantic 代码塞进外部库类本身。

用 GetPydanticSchema 减少样板代码

若只需要简单的转换,通常不必定义标记类。Pydantic 提供 GetPydanticSchema,可把钩子逻辑直接放进 Annotated。以下例子让字符串在验证后重复一次:

from typing import Annotated

from pydantic_core import core_schema

from pydantic import BaseModel, GetPydanticSchema

class Model(BaseModel):
    y: Annotated[
        str,
        GetPydanticSchema(
            lambda tp, handler: core_schema.no_info_after_validator_function(
                lambda x: x * 2, handler(tp)
            )
        ),
    ]

assert Model(y='ab').y == 'abab'

支持自定义泛型类

在字段中使用 Python 泛型类时,可以通过 __get_pydantic_core_schema__ 按其类型参数验证成员。这是进阶方法;多数场景下,普通 Pydantic 模型已经足够。

泛型类若定义了类方法 __get_pydantic_core_schema__,不必为它设置 arbitrary_types_allowed。由于 source_type 与 cls 不同,可用 typing.get_args(或 typing_extensions.get_args)提取子类型;再用 handler.generate_schema() 为子类型生成独立 Schema。不要简单地用 handler(item_type),因为泛型参数需要独立于当前 Annotated 元数据上下文生成 Schema。

以下 Owner 泛型保存一个名称和一个泛型 item。它先取得 item 类型 Schema,确保传入值为 Owner 实例,再用包裹验证器将 item 交给子类型 Schema。JSON 输入则先验证具有 name 和 item 字段的对象,再转换为 Owner 实例:

from dataclasses import dataclass
from typing import Any, Generic, TypeVar

from pydantic_core import CoreSchema, core_schema
from typing_extensions import get_args, get_origin

from pydantic import (
    BaseModel,
    GetCoreSchemaHandler,
    ValidatorFunctionWrapHandler,
)

ItemType = TypeVar('ItemType')

@dataclass
class Owner(Generic[ItemType]):
    name: str
    item: ItemType

    @classmethod
    def __get_pydantic_core_schema__(
        cls, source_type: Any, handler: GetCoreSchemaHandler
    ) -> CoreSchema:
        origin = get_origin(source_type)
        if origin is None:
            origin = source_type
            item_type = Any
        else:
            item_type = get_args(source_type)[0]

        item_schema = handler.generate_schema(item_type)

        def validate_item(
            value: Owner[Any], wrap_handler: ValidatorFunctionWrapHandler
        ) -> Owner[Any]:
            value.item = wrap_handler(value.item)
            return value

        python_schema = core_schema.chain_schema(
            [
                core_schema.is_instance_schema(cls),
                core_schema.no_info_wrap_validator_function(
                    validate_item, item_schema
                ),
            ]
        )

        return core_schema.json_or_python_schema(
            json_schema=core_schema.chain_schema(
                [
                    core_schema.typed_dict_schema(
                        {
                            'name': core_schema.typed_dict_field(
                                core_schema.str_schema()
                            ),
                            'item': core_schema.typed_dict_field(item_schema),
                        }
                    ),
                    core_schema.no_info_before_validator_function(
                        lambda data: Owner(
                            name=data['name'], item=data['item']
                        ),
                        python_schema,
                    ),
                ]
            ),
            python_schema=python_schema,
        )

class Car(BaseModel):
    color: str

class House(BaseModel):
    rooms: int

class Model(BaseModel):
    car_owner: Owner[Car]
    home_owner: Owner[House]

model = Model(
    car_owner=Owner(name='John', item=Car(color='black')),
    home_owner=Owner(name='James', item=House(rooms=3)),
)

model_from_json = Model.model_validate_json(
    '{"car_owner":{"name":"John","item":{"color":"black"}},'
    '"home_owner":{"name":"James","item":{"rooms":3}}}'
)

无类型参数时,示例把 item 类型视为 Any;有参数时,则提取对应子类型并调用 generate_schema()。页面还展示了把 Car 和 House 类型对调后触发验证错误,以及 JSON 输入中字段结构不符时报告具体成员错误。此处仅说明错误路径,未把页面输出中与示例代码不对应的错误文本照录。

构造自定义泛型序列

相同方式可用于自定义容器。下面的 MySequence 实现 Sequence 接口;有元素类型参数时,让 Pydantic 生成对应 Sequence 子类型 Schema,然后把经验证的序列封装为 MySequence。

from collections.abc import Sequence
from typing import Any, TypeVar

from pydantic_core import ValidationError, core_schema
from typing_extensions import get_args

from pydantic import BaseModel, GetCoreSchemaHandler

T = TypeVar('T')

class MySequence(Sequence[T]):
    def __init__(self, value: Sequence[T]):
        self.v = value

    def __getitem__(self, index):
        return self.v[index]

    def __len__(self):
        return len(self.v)

    @classmethod
    def __get_pydantic_core_schema__(
        cls, source: Any, handler: GetCoreSchemaHandler
    ) -> core_schema.CoreSchema:
        instance_schema = core_schema.is_instance_schema(cls)

        args = get_args(source)
        if args:
            sequence_schema = handler.generate_schema(Sequence[args[0]])
        else:
            sequence_schema = handler.generate_schema(Sequence)

        non_instance_schema = core_schema.no_info_after_validator_function(
            MySequence, sequence_schema
        )
        return core_schema.union_schema([instance_schema, non_instance_schema])

class M(BaseModel):
    model_config = dict(validate_default=True)

    s1: MySequence = [3]

m = M()
print(m.s1.v)
# [3]

class TypedModel(BaseModel):
    s1: MySequence[int]

TypedModel(s1=[1])
try:
    TypedModel(s1=['a'])
except ValidationError as exc:
    print(exc)

这里默认值验证通过 model_config 中的 validate_default=True 开启。未指定泛型参数时,生成一般 Sequence Schema;指定 MySequence[int] 时,会按整数序列验证,不能解析为整数的元素会产生验证错误。

在验证器中读取字段名

Pydantic V2.4 起,可以在 __get_pydantic_core_schema__ 中读取 handler.field_name,并让验证函数从 ValidationInfo.field_name 获得字段名。该能力在 V2.4 重新加入;V2.0 至 V2.3 不支持。

from typing import Any

from pydantic_core import core_schema

from pydantic import BaseModel, GetCoreSchemaHandler, ValidationInfo

class CustomType:
    def __init__(self, value: int, field_name: str):
        self.value = value
        self.field_name = field_name

    def __repr__(self):
        return f'CustomType<{self.value} {self.field_name!r}>'

    @classmethod
    def validate(cls, value: int, info: ValidationInfo):
        return cls(value, info.field_name)

    @classmethod
    def __get_pydantic_core_schema__(
        cls, source_type: Any, handler: GetCoreSchemaHandler
    ) -> core_schema.CoreSchema:
        return core_schema.with_info_after_validator_function(
            cls.validate, handler(int)
        )

class MyModel(BaseModel):
    my_field: CustomType

model = MyModel(my_field=1)
print(model.my_field)
# CustomType<1 'my_field'>

Annotated 元数据也可以从 ValidationInfo 读取字段名,例如 AfterValidator 接收 info 参数:

from typing import Annotated

from pydantic import AfterValidator, BaseModel, ValidationInfo

def my_validator(value: int, info: ValidationInfo):
    return f'<{value} {info.field_name!r}>'

class MyModel(BaseModel):
    my_field: Annotated[int, AfterValidator(my_validator)]

model = MyModel(my_field=1)
print(model.my_field)
# <1 'my_field'>

选择适合的集成层

可以按所需控制范围逐步选择实现方式:

  1. 先使用标准 Python 类型以及 Pydantic 的 Field、Annotated 和验证器标记;
  2. 若需要简洁地自定义底层验证,用 GetPydanticSchema;
  3. 若需要可复用的复杂适配逻辑,在 Annotated 元数据标记类中实现 __get_pydantic_core_schema__;
  4. 若该行为本来就属于自己的类型,也可把钩子直接实现到该类上;
  5. 若目标类型来自第三方库且不能修改,用 Annotated 包装它,并显式定义验证、序列化和 JSON Schema;
  6. 自定义泛型需要根据 source_type 提取参数,并使用 handler.generate_schema() 生成其成员 Schema;
  7. 对递归定义使用命名类型别名;需要字段名时,Pydantic V2.4 及以上可读取 ValidationInfo.field_name。

底层 CoreSchema 钩子提供了精细控制,但也使实现依赖更低层的 API。对于简单约束和变换,高层的 Annotated、Field 与内置验证器通常更容易维护。

来源信息

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容