为 Python dataclass 加上 Pydantic 输入校验

本文完整译编自 Pydantic 官方文档 Dataclasses,作者为 Pydantic 文档贡献者。示例针对 Pydantic 2 系列 API;来源是会更新的 latest 文档,本稿核验于 2026 年 10 月 5 日,未运行示例。代码输出均为原文展示的预期输出。

在 dataclass 中使用输入校验

如果不想使用 Pydantic 的 BaseModel,可以让标准数据类获得同样的输入校验能力。把装饰器换成 pydantic.dataclasses.dataclass:

from datetime import datetime
from typing import Optional

from pydantic.dataclasses import dataclass


@dataclass
class User:
    id: int
    name: str = 'John Doe'
    signup_ts: Optional[datetime] = None


user = User(id='42', signup_ts='2032-06-21T12:00')
print(user)
"""
User(id=42, name='John Doe', signup_ts=datetime.datetime(2032, 6, 21, 12, 0))
"""

这个例子把字符串 "42" 转为整数,也把日期时间字符串转为 datetime 对象。API 入口是 @pydantic.dataclasses.dataclass。

请记住,Pydantic dataclass 并不是 Pydantic 模型的替代品。它保持了类似标准库 dataclass 的功能,再加上 Pydantic 校验;有些场景仍更适合 BaseModel。关于两者定位的讨论,参见 pydantic/pydantic#710。

与 Pydantic 模型的相同点和不同点

两者都支持配置和嵌套类型,实例化时传入的参数也会被复制。但 dataclass 不使用模型的 model_config 属性;它的配置入口见后文。

dataclass 不提供模型上用于校验、导出和生成 JSON Schema 的整套方法。需要这些能力时,把数据类交给 TypeAdapter,再使用适配器的方法:

from pydantic import TypeAdapter
from pydantic.dataclasses import dataclass


@dataclass
class Foo:
    f: int


foo = Foo(f=1)

TypeAdapter(Foo).dump_python(foo)
#> {'f': 1}
TypeAdapter(Foo).validate_python({'f': 1})
#> Foo(f=1)

自定义校验器同样可用,但初始化相关行为有所不同,本文最后一节给出完整示例。额外字段的行为也有区别:额外数据不会包含在序列化结果中,而且不能通过 __pydantic_extra__ 属性定制额外值的校验。使用 extra 配置时,不能从对象是否接收某个值推断它会出现在导出结果中;标准 dataclass 的表示也主要围绕已声明字段。

泛型数据类的一个重要限制

Pydantic 支持泛型 dataclass,但直接实例化带类型参数的别名,并不能像泛型 BaseModel 那样保证对类型参数进行校验。原文给出两种 Python 语法。Python 3.9 及以上的传统写法是:

from typing import Generic, TypeVar

from pydantic.dataclasses import dataclass

T = TypeVar('T')


@dataclass
class Foo(Generic[T]):
    f: T


Foo[int](f='not_an_int')  # (1)
#> Foo(f='not_an_int')

在这里,Foo[int] 是泛型别名,不是一个真正独立的类型对象。Pydantic 当前把这种构造视为 Foo[Any],因此 f 没有按整数要求校验,字符串照样留在结果里。

Python 3.12 及以上可以使用新的类型参数语法,同样存在这个限制:

from pydantic.dataclasses import dataclass


@dataclass
class Foo[T]:
    f: T


Foo[int](f='not_an_int')  # (1)
#> Foo(f='not_an_int')

需要按具体类型参数校验时,应把 Foo[int] 交给 TypeAdapter,再使用适配器执行校验,例如通过 TypeAdapter(Foo[int]).validate_python(...) 这一入口。上述 Python 版本标签说明示例语法,不替代你所安装 Pydantic 版本的 Python 支持范围。

同时使用 Field 和标准库 field

可以在同一个数据类中同时使用 Pydantic 的 Field() 和标准库的 dataclasses.field():

import dataclasses
from typing import Optional

from pydantic import Field
from pydantic.dataclasses import dataclass


@dataclass
class User:
    id: int
    name: str = 'John Doe'
    friends: list[int] = dataclasses.field(default_factory=lambda: [0])
    age: Optional[int] = dataclasses.field(
        default=None,
        metadata={'title': 'The age of the user', 'description': 'do not lie!'},
    )
    height: Optional[int] = Field(
        default=None, title='The height in cm', ge=50, le=300
    )


user = User(id='42', height='250')
print(user)
#> User(id=42, name='John Doe', friends=[0], age=None, height=250)

这个例子用 default_factory 提供 friends 的默认列表,用标准库字段 metadata 提供 age 的标题和描述,再用 Field 将 height 限制在 50 到 300 之间。输入字符串形式的 height 会进入整数校验流程。Pydantic 的 @dataclass 接收与标准装饰器相同的参数,另外增加 config。

配置 dataclass

若要像配置 BaseModel 那样修改校验行为,有两个入口:在装饰器中传入 config,或定义 __pydantic_config__ 属性。下面两种写法都启用赋值校验:

from pydantic import ConfigDict
from pydantic.dataclasses import dataclass


# Option 1 -- using the decorator argument:
@dataclass(config=ConfigDict(validate_assignment=True))  # (1)
class MyDataclass1:
    a: int


# Option 2 -- using an attribute:
@dataclass
class MyDataclass2:
    a: int

    __pydantic_config__ = ConfigDict(validate_assignment=True)

这里的 validate_assignment=True 用于属性赋值时的校验;配置详情见 validate_assignment 参考文档。不要把构造时校验等同于以后所有变异都会自动重验,尤其是嵌套对象或容器内部的修改。

重建 dataclass 的 schema

rebuild_dataclass() 可以重建数据类的核心 schema。关于何时需要重建以及前向引用等问题,参见 函数参考 和 重建模型 schema 说明。

标准库 dataclass 与 Pydantic dataclass

继承标准库数据类

Pydantic 数据类可以继承标准库数据类,包括存在嵌套的情况。继承得到的字段也会自动校验:

import dataclasses

import pydantic


@dataclasses.dataclass
class Z:
    z: int


@dataclasses.dataclass
class Y(Z):
    y: int = 0


@pydantic.dataclasses.dataclass
class X(Y):
    x: int = 0


foo = X(x=b'1', y='2', z='3')
print(foo)
#> X(z=3, y=2, x=1)

try:
    X(z='pika')
except pydantic.ValidationError as e:
    print(e)
    """
    1 validation error for X
    z
      Input should be a valid integer, unable to parse string as an integer [type=int_parsing, input_value='pika', input_type=str]
    """

示例中 z、y、x 均被转换为整数,而 z="pika" 会产生定位到 z 字段的 int_parsing 校验错误。Pydantic dataclass 像模型一样校验输入,因此也具备相同的可观测性入口:如果使用 Logfire,数据类校验会与模型校验一起记录,结构化错误中包含被拒绝的值。编者提示:接入日志或可观测性服务时,应据实际数据检查敏感输入的记录策略。

也可以直接把 Pydantic 装饰器应用到一个已有的标准库 dataclass 上;这会创建一个新的子类:

import dataclasses

import pydantic


@dataclasses.dataclass
class A:
    a: int


PydanticA = pydantic.dataclasses.dataclass(A)
print(PydanticA(a='1'))
#> A(a=1)

在 BaseModel 内使用标准库数据类

当标准库数据类作为 Pydantic 模型、Pydantic 数据类或 TypeAdapter 的一部分使用时,也会应用校验,同时保留该数据类的配置。因此,把标准库数据类或 Pydantic 数据类写成字段类型注解,在校验能力上是等价的。

import dataclasses
from typing import Optional

from pydantic import BaseModel, ConfigDict, ValidationError


@dataclasses.dataclass(frozen=True)
class User:
    name: str


class Foo(BaseModel):
    # Required so that pydantic revalidates the model attributes:
    model_config = ConfigDict(revalidate_instances='always')

    user: Optional[User] = None


# nothing is validated as expected:
user = User(name=['not', 'a', 'string'])
print(user)
#> User(name=['not', 'a', 'string'])


try:
    Foo(user=user)
except ValidationError as e:
    print(e)
    """
    1 validation error for Foo
    user.name
      Input should be a valid string [type=string_type, input_value=['not', 'a', 'string'], input_type=list]
    """

foo = Foo(user=User(name='pika'))
try:
    foo.user.name = 'bulbi'
except dataclasses.FrozenInstanceError as e:
    print(e)
    #> cannot assign to field 'name'

这里直接构造标准库的 User 不会校验 name,所以列表暂时被接受。交给 Foo 时,由于显式设置了 revalidate_instances="always",已有实例会重新校验,错误定位为 user.name。合法对象构造后,对 foo.user.name 重新赋值又会受到标准库 frozen=True 限制,抛出 FrozenInstanceError。这两个行为分别来自 Pydantic 的重验配置和 dataclass 的冻结配置。

使用自定义类型

由于标准库数据类进入 Pydantic 上下文后也会校验,如果字段使用了 Pydantic 不认识的自定义类型,就可能在引用该数据类时遇到 schema 生成错误。可以配置 arbitrary_types_allowed 来允许此类类型:

import dataclasses

from pydantic import BaseModel, ConfigDict
from pydantic.errors import PydanticSchemaGenerationError


class ArbitraryType:
    def __init__(self, value):
        self.value = value

    def __repr__(self):
        return f'ArbitraryType(value={self.value!r})'


@dataclasses.dataclass
class DC:
    a: ArbitraryType
    b: str


# valid as it is a stdlib dataclass without validation:
my_dc = DC(a=ArbitraryType(value=3), b='qwe')

try:

    class Model(BaseModel):
        dc: DC
        other: str

    # invalid as dc is now validated with pydantic, and ArbitraryType is not a known type
    Model(dc=my_dc, other='other')

except PydanticSchemaGenerationError as e:
    print(e.message)
    """
    Unable to generate pydantic-core schema for <class '__main__.ArbitraryType'>. Set `arbitrary_types_allowed=True` in the model_config to ignore this error or implement `__get_pydantic_core_schema__` on your type to fully support it.

    If you got this error by calling handler(<some type>) within `__get_pydantic_core_schema__` then you likely need to call `handler.generate_schema(<some type>)` since we do not call `__get_pydantic_core_schema__` on `<some type>` otherwise to avoid infinite recursion.
    """


# valid as we set arbitrary_types_allowed=True, and that config pushes down to the nested vanilla dataclass
class Model(BaseModel):
    model_config = ConfigDict(arbitrary_types_allowed=True)

    dc: DC
    other: str


m = Model(dc=my_dc, other='other')
print(repr(m))
#> Model(dc=DC(a=ArbitraryType(value=3), b='qwe'), other='other')

这段示例先直接构造标准库 DC,此时没有 Pydantic 校验;随后在 Model 中引用 DC,由于 ArbitraryType 未知而产生 PydanticSchemaGenerationError。错误信息给出两个方向:允许任意类型,或在类型上实现 __get_pydantic_core_schema__ 以完整支持它。若在该方法里调用 handler 导致相关问题,应按错误提示考虑 handler.generate_schema(...),避免递归处理错误。

最终示例把 arbitrary_types_allowed=True 放在外层模型的 model_config 中,配置会传递到嵌套的普通数据类。编者注:这项配置主要允许并检查相应的自定义类型实例,不会自动证明实例内部的 value 合法;若内部状态也有业务约束,需要为它建立明确的校验。

判断一个类是不是 Pydantic dataclass

Pydantic 数据类仍然是 dataclass,因此 dataclasses.is_dataclass() 会返回 True。要专门判断它是否经过 Pydantic 包装,应使用 pydantic.dataclasses.is_pydantic_dataclass():

import dataclasses

import pydantic


@dataclasses.dataclass
class StdLibDataclass:
    id: int


PydanticDataclass = pydantic.dataclasses.dataclass(StdLibDataclass)

print(dataclasses.is_dataclass(StdLibDataclass))
#> True
print(pydantic.dataclasses.is_pydantic_dataclass(StdLibDataclass))
#> False

print(dataclasses.is_dataclass(PydanticDataclass))
#> True
print(pydantic.dataclasses.is_pydantic_dataclass(PydanticDataclass))
#> True

校验器与初始化钩子

字段校验器可以用于 Pydantic 数据类。下面的 before 校验器在字符串校验前把整数转成字符串,并在左侧补零:

from pydantic import field_validator
from pydantic.dataclasses import dataclass


@dataclass
class DemoDataclass:
    product_id: str  # should be a five-digit string, may have leading zeros

    @field_validator('product_id', mode='before')
    @classmethod
    def convert_int_serial(cls, v):
        if isinstance(v, int):
            v = str(v).zfill(5)
        return v


print(DemoDataclass(product_id='01234'))
#> DemoDataclass(product_id='01234')
print(DemoDataclass(product_id=2468))
#> DemoDataclass(product_id='02468')

编者注:原注释说 product_id 应为五位数字字符串,但这段示例只演示转换,未完整实现这一业务规则。zfill(5) 只保证最小宽度,不会拒绝负数、超长整数或非数字字符串;另外 Python 的 bool 也是 int 的子类。如业务确实要求恰好五位十进制数字,应再明确长度、字符集、数值范围及布尔值策略。本文保留原示例,避免把转换函数误写成完整验证器。

dataclass 的 __post_init__() 也受支持;它位于 before 和 after 两类模型校验器之间调用:

Pydantic dataclass:校验与初始化的先后顺序
图 1:原创初始化流程示意。下例的输出顺序为 First、Second、Third。
from pydantic_core import ArgsKwargs
from typing_extensions import Self

from pydantic import model_validator
from pydantic.dataclasses import dataclass


@dataclass
class Birth:
    year: int
    month: int
    day: int


@dataclass
class User:
    birth: Birth

    @model_validator(mode='before')
    @classmethod
    def before(cls, values: ArgsKwargs) -> ArgsKwargs:
        print(f'First: {values}')  # (1)
        """
        First: ArgsKwargs((), {'birth': {'year': 1995, 'month': 3, 'day': 2}})
        """
        return values

    @model_validator(mode='after')
    def after(self) -> Self:
        print(f'Third: {self}')
        #> Third: User(birth=Birth(year=1995, month=3, day=2))
        return self

    def __post_init__(self):
        print(f'Second: {self.birth}')
        #> Second: Birth(year=1995, month=3, day=2)


user = User(**{'birth': {'year': 1995, 'month': 3, 'day': 2}})

与 Pydantic 模型的对应入口不同,示例中 before 校验器收到的 values 是 ArgsKwargs,不是可以一概按字典处理的对象。before 返回参数后,birth 的字典进入嵌套 Birth 数据类的校验和构造;__post_init__ 因此看到已经构造的 Birth;最后 after 接收 self 并返回它。

编者注:Birth 示例仅声明 year、month、day 为整数,没有校验真实日期是否存在。初始化示例使用 print 展示输入及对象,若应用于真实个人信息,应避免直接照搬到生产日志。

来源与许可证

来源:Pydantic 官方 Dataclasses 文档,页面标注 © Pydantic Services Inc. 2025 to present。示例代码及相关文档保留 Pydantic 仓库的 MIT 许可证 通知如下。中文译文、技术边界说明与原创图由未完纪经授权整理;没有将原文输出称作本次运行结果。

The MIT License (MIT)

Copyright (c) 2017 to present Pydantic Services Inc. and individual contributors.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容