MongoDB 聚合:对订单分组并汇总

本教程介绍如何构建聚合管道,对集合执行聚合,并使用你选择的语言显示结果。

任务说明

本例对客户订单数据进行分组与分析。结果列出在 2020 年购买过商品的客户,并包含每位客户在 2020 年的订单历史。

聚合管道执行以下操作:

  • 按字段值匹配文档子集。
  • 将具有相同字段值的文档分组。
  • 为每个结果文档添加计算字段。

开始之前

在 MongoDB 官方页面中,使用右上角的“Select your language”下拉菜单选择示例语言,或选择 MongoDB Shell。下文按语言展开各分支的准备工作与操作步骤。

示例使用 orders 集合保存单个商品订单。每个订单只对应一位客户,因此按保存客户邮箱的 customer_id 字段分组;C# 模型对应的属性名为 CustomerId。

使用专门的练习数据库准备样本。部分原例中的 deleteMany、DeleteMany 或 delete_many 会清空目标集合;其他分支直接追加数据,重复执行插入会重复计入订单。

MongoDB Shell

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

使用 db.collection.insertMany() 创建 orders 集合并插入以下样本:

db.orders.insertMany( [
   {
      customer_id: "elise_smith@myemail.com",
      orderdate: new Date("2020-05-30T08:35:52Z"),
      value: 231,
   },
   {
      customer_id: "elise_smith@myemail.com",
      orderdate: new Date("2020-01-13T09:32:07Z"),
      value: 99,
   },
   {
      customer_id: "oranieri@warmmail.com",
      orderdate: new Date("2020-01-01T08:25:37Z"),
      value: 63,
   },
   {
      customer_id: "tj@wheresmyemail.com",
      orderdate: new Date("2019-05-28T19:13:32Z"),
      value: 2,
   },
   {
      customer_id: "tj@wheresmyemail.com",
      orderdate: new Date("2020-11-23T22:56:53Z"),
      value: 187,
   },
   {
      customer_id: "tj@wheresmyemail.com",
      orderdate: new Date("2020-08-18T23:04:48Z"),
      value: 4,
   },
   {
      customer_id: "elise_smith@myemail.com",
      orderdate: new Date("2020-12-26T08:55:46Z"),
      value: 4,
   },
   {
      customer_id: "tj@wheresmyemail.com",
      orderdate: new Date("2021-02-29T07:49:32Z"),
      value: 1024,
   },
   {
      customer_id: "elise_smith@myemail.com",
      orderdate: new Date("2020-10-03T13:49:44Z"),
      value: 102,
   }
] )

操作步骤

1. 运行聚合管道

db.orders.aggregate( [
   // Stage 1: Match orders in 2020
   { $match: {
      orderdate: {
         $gte: new Date("2020-01-01T00:00:00Z"),
         $lt: new Date("2021-01-01T00:00:00Z"),
      }
   } },

   // Stage 2: Sort orders by date
   { $sort: { orderdate: 1 } },

   // Stage 3: Group orders by email address (customer_id)
   { $group: {
      _id: "$customer_id",
      first_purchase_date: { $first: "$orderdate" },
      total_value: { $sum: "$value" },
      total_orders: { $sum: 1 },
      orders: { $push:
         {
            orderdate: "$orderdate",
            value: "$value"
         }
      }
   } },

   // Stage 4: Sort orders by first order date
   { $sort: { first_purchase_date: 1 } },

   // Stage 5: Display the customers' email addresses
   { $set: { customer_id: "$_id" } },

   // Stage 6: Remove unneeded fields
   { $unset: ["_id"] }
] )

2. 解读聚合结果

聚合返回 2020 年客户订单的以下汇总。结果按客户邮箱分组,包含这位客户在该年下单的全部订单详情。

{
   first_purchase_date: ISODate("2020-01-01T08:25:37.000Z"),
   total_value: 63,
   total_orders: 1,
   orders: [ { orderdate: ISODate("2020-01-01T08:25:37.000Z"), value: 63 } ],
   customer_id: 'oranieri@warmmail.com'
}
{
   first_purchase_date: ISODate("2020-01-13T09:32:07.000Z"),
   total_value: 436,
   total_orders: 4,
   orders: [
      { orderdate: ISODate("2020-01-13T09:32:07.000Z"), value: 99 },
      { orderdate: ISODate("2020-05-30T08:35:52.000Z"), value: 231 },
      { orderdate: ISODate("2020-10-03T13:49:44.000Z"), value: 102 },
      { orderdate: ISODate("2020-12-26T08:55:46.000Z"), value: 4 }
   ],
   customer_id: 'elise_smith@myemail.com'
}
{
   first_purchase_date: ISODate("2020-08-18T23:04:48.000Z"),
   total_value: 191,
   total_orders: 2,
   orders: [
      { orderdate: ISODate("2020-08-18T23:04:48.000Z"), value: 4 },
      { orderdate: ISODate("2020-11-23T22:56:53.000Z"), value: 187 }
   ],
   customer_id: 'tj@wheresmyemail.com'
}

C

创建模板应用

开始本教程之前,先创建一个新的 C 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg-tutorial.c 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

#include <stdio.h>
#include <bson/bson.h>
#include <mongoc/mongoc.h>

int main(void)
{
    mongoc_init();

    // Replace the placeholder with your connection string.
    char *uri = "<connection string>";
    mongoc_client_t* client = mongoc_client_new(uri);

    // Get a reference to relevant collections.
    // ... mongoc_collection_t *some_coll = mongoc_client_get_collection(client, "agg_tutorials_db", "some_coll");
    // ... mongoc_collection_t *another_coll = mongoc_client_get_collection(client, "agg_tutorials_db", "another_coll");

    // Delete any existing documents in collections if needed.
    // ... {
    // ...     bson_t *filter = bson_new();
    // ...     bson_error_t error;
    // ...     if (!mongoc_collection_delete_many(some_coll, filter, NULL, NULL, &error))
    // ...     {
    // ...         fprintf(stderr, "Delete error: %s\n", error.message);
    // ...     }
    // ...     bson_destroy(filter);
    // ... }

    // Insert sample data into the collection or collections.
    // ... {
    // ...     size_t num_docs = ...;
    // ...     bson_t *docs[num_docs];
    // ...
    // ...     docs[0] = ...;
    // ...
    // ...     bson_error_t error;
    // ...     if (!mongoc_collection_insert_many(some_coll, (const bson_t **)docs, num_docs, NULL, NULL, &error))
    // ...     {
    // ...         fprintf(stderr, "Insert error: %s\n", error.message);
    // ...     }
    // ...
    // ...     for (int i = 0; i < num_docs; i++)
    // ...     {
    // ...         bson_destroy(docs[i]);
    // ...     }
    // ... }

    {
        const bson_t *doc;

        // Add code to create pipeline stages.
        bson_t *pipeline = BCON_NEW("pipeline", "[",
        // ... Add pipeline stages here.
        "]");

        // Run the aggregation.
        // ... mongoc_cursor_t *results = mongoc_collection_aggregate(some_coll, MONGOC_QUERY_NONE, pipeline, NULL, NULL);

        bson_destroy(pipeline);

        // Print the aggregation results.
        while (mongoc_cursor_next(results, &doc))
        {
            char *str = bson_as_canonical_extended_json(doc, NULL);
            printf("%s\n", str);
            bson_free(str);
        }
        bson_error_t error;
        if (mongoc_cursor_error(results, &error))
        {
            fprintf(stderr, "Aggregation error: %s\n", error.message);
        }

        mongoc_cursor_destroy(results);
    }

    // Clean up resources.
    // ... mongoc_collection_destroy(some_coll);
    mongoc_client_destroy(client);
    mongoc_cleanup();

    return EXIT_SUCCESS;
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

char *uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

mongoc_collection_t *orders = mongoc_client_get_collection(client, "agg_tutorials_db", "orders");

{
    bson_t *filter = bson_new();
    bson_error_t error;
    if (!mongoc_collection_delete_many(orders, filter, NULL, NULL, &error))
    {
        fprintf(stderr, "Delete error: %s\n", error.message);
    }
    bson_destroy(filter);
}

{
    size_t num_docs = 9;
    bson_t *docs[num_docs];

    docs[0] = BCON_NEW(
        "customer_id", BCON_UTF8("elise_smith@myemail.com"),
        "orderdate", BCON_DATE_TIME(1590825352000UL), // 2020-05-30T08:35:52Z
        "value", BCON_INT32(231));

    docs[1] = BCON_NEW(
        "customer_id", BCON_UTF8("elise_smith@myemail.com"),
        "orderdate", BCON_DATE_TIME(1578904327000UL), // 2020-01-13T09:32:07Z
        "value", BCON_INT32(99));

    docs[2] = BCON_NEW(
        "customer_id", BCON_UTF8("oranieri@warmmail.com"),
        "orderdate", BCON_DATE_TIME(1577865937000UL), // 2020-01-01T08:25:37Z
        "value", BCON_INT32(63));

    docs[3] = BCON_NEW(
        "customer_id", BCON_UTF8("tj@wheresmyemail.com"),
        "orderdate", BCON_DATE_TIME(1559061212000UL), // 2019-05-28T19:13:32Z
        "value", BCON_INT32(2));

    docs[4] = BCON_NEW(
        "customer_id", BCON_UTF8("tj@wheresmyemail.com"),
        "orderdate", BCON_DATE_TIME(1606171013000UL), // 2020-11-23T22:56:53Z
        "value", BCON_INT32(187));

    docs[5] = BCON_NEW(
        "customer_id", BCON_UTF8("tj@wheresmyemail.com"),
        "orderdate", BCON_DATE_TIME(1597793088000UL), // 2020-08-18T23:04:48Z
        "value", BCON_INT32(4));

    docs[6] = BCON_NEW(
        "customer_id", BCON_UTF8("elise_smith@myemail.com"),
        "orderdate", BCON_DATE_TIME(1608963346000UL), // 2020-12-26T08:55:46Z
        "value", BCON_INT32(4));

    docs[7] = BCON_NEW(
        "customer_id", BCON_UTF8("tj@wheresmyemail.com"),
        "orderdate", BCON_DATE_TIME(1614496172000UL), // 2021-02-28T07:49:32Z
        "value", BCON_INT32(1024));

    docs[8] = BCON_NEW(
        "customer_id", BCON_UTF8("elise_smith@myemail.com"),
        "orderdate", BCON_DATE_TIME(1601722184000UL), // 2020-10-03T13:49:44Z
        "value", BCON_INT32(102));

    bson_error_t error;
    if (!mongoc_collection_insert_many(orders, (const bson_t **)docs, num_docs, NULL, NULL, &error))
    {
        fprintf(stderr, "Insert error: %s\n", error.message);
    }

    for (int i = 0; i < num_docs; i++)
    {
        bson_destroy(docs[i]);
    }
}

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

"{", "$match", "{",
"orderdate", "{",
"$gte", BCON_DATE_TIME(1577836800000UL), // Represents 2020-01-01T00:00:00Z
"$lt", BCON_DATE_TIME(1609459200000UL),  // Represents 2021-01-01T00:00:00Z
"}",
"}", "}",

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

"{", "$sort", "{", "orderdate", BCON_INT32(1), "}", "}",

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
"{", "$group", "{",
"_id", BCON_UTF8("$customer_id"),
"first_purchase_date", "{", "$first", BCON_UTF8("$orderdate"), "}",
"total_value", "{", "$sum", BCON_UTF8("$value"), "}",
"total_orders", "{", "$sum", BCON_INT32(1), "}",
"orders", "{", "$push", "{",
"orderdate", BCON_UTF8("$orderdate"),
"value", BCON_UTF8("$value"),
"}", "}",
"}", "}",

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

"{", "$sort", "{", "first_purchase_date", BCON_INT32(1), "}", "}",

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

"{", "$set", "{", "customer_id", BCON_UTF8("$_id"), "}", "}",

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

"{", "$unset", "[", BCON_UTF8("_id"), "]", "}",

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

mongoc_cursor_t *results =
    mongoc_collection_aggregate(orders, MONGOC_QUERY_NONE, pipeline, NULL, NULL);

bson_destroy(pipeline);

在清理语句中加入以下一行,释放集合资源:

mongoc_collection_destroy(orders);

最后在终端中执行以下命令,生成并运行可执行文件:

gcc -o aggc agg-tutorial.c $(pkg-config --libs --cflags libmongoc-1.0)
./aggc

如果一次调用上述命令时遇到连接错误,可以把编译与执行分开运行。

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{ "first_purchase_date" : { "$date" : { "$numberLong" : "1577865937000" } }, "total_value" : { "$numberInt" : "63" }, "total_orders" : { "$numberInt" : "1" }, "orders" : [ { "orderdate" : { "$date" : { "$numberLong" : "1577865937000" } }, "value" : { "$numberInt" : "63" } } ], "customer_id" : "oranieri@warmmail.com" }
{ "first_purchase_date" : { "$date" : { "$numberLong" : "1578904327000" } }, "total_value" : { "$numberInt" : "436" }, "total_orders" : { "$numberInt" : "4" }, "orders" : [ { "orderdate" : { "$date" : { "$numberLong" : "1578904327000" } }, "value" : { "$numberInt" : "99" } }, { "orderdate" : { "$date" : { "$numberLong" : "1590825352000" } }, "value" : { "$numberInt" : "231" } }, { "orderdate" : { "$date" : { "$numberLong" : "1601722184000" } }, "value" : { "$numberInt" : "102" } }, { "orderdate" : { "$date" : { "$numberLong" : "1608963346000" } }, "value" : { "$numberInt" : "4" } } ], "customer_id" : "elise_smith@myemail.com" }
{ "first_purchase_date" : { "$date" : { "$numberLong" : "1597793088000" } }, "total_value" : { "$numberInt" : "191" }, "total_orders" : { "$numberInt" : "2" }, "orders" : [ { "orderdate" : { "$date" : { "$numberLong" : "1597793088000" } }, "value" : { "$numberInt" : "4" } }, { "orderdate" : { "$date" : { "$numberLong" : "1606171013000" } }, "value" : { "$numberInt" : "187" } } ], "customer_id" : "tj@wheresmyemail.com" }

结果按客户邮箱分组,包含这位客户的全部订单详情。

C++

创建模板应用

开始本教程之前,先创建一个新的 C++ 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

C++ 驱动的更多用法可参阅API 文档。

安装驱动或库后,创建 agg-tutorial.cpp 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

#include <iostream>
#include <bsoncxx/builder/basic/document.hpp>
#include <bsoncxx/builder/basic/kvp.hpp>
#include <bsoncxx/json.hpp>
#include <mongocxx/client.hpp>
#include <mongocxx/instance.hpp>
#include <mongocxx/pipeline.hpp>
#include <mongocxx/uri.hpp>
#include <chrono>

using bsoncxx::builder::basic::kvp;
using bsoncxx::builder::basic::make_document;
using bsoncxx::builder::basic::make_array;

int main() {
   mongocxx::instance instance;

   // Replace the placeholder with your connection string.
   mongocxx::uri uri("<connection string>");
   mongocxx::client client(uri);

   auto db = client["agg_tutorials_db"];
   // Delete existing data in the database, if necessary.
   db.drop();

   // Get a reference to relevant collections.
   // ... auto some_coll = db["..."];
   // ... auto another_coll = db["..."];

   // Insert sample data into the collection or collections.
   // ... some_coll.insert_many(docs);

   // Create an empty pipelne.
   mongocxx::pipeline pipeline;

   // Add code to create pipeline stages.
   // pipeline.match(make_document(...));

   // Run the aggregation and print the results.
   auto cursor = orders.aggregate(pipeline);
   for (auto&& doc : cursor) {
      std::cout << bsoncxx::to_json(doc, bsoncxx::ExtendedJsonMode::k_relaxed) << std::endl;
   }
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

mongocxx::uri uri{"mongodb+srv://mongodb-example:27017"};

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

auto orders = db["orders"];

std::vector<bsoncxx::document::value> docs = {
    bsoncxx::from_json(R"({
        "customer_id": "elise_smith@myemail.com",
        "orderdate": {"$date": 1590821752000},
        "value": 231
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "elise_smith@myemail.com",
        "orderdate": {"$date": 1578901927},
        "value": 99
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "oranieri@warmmail.com",
        "orderdate": {"$date": 1577861137},
        "value": 63
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": {"$date": 1559076812},
        "value": 2
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": {"$date": 1606172213},
        "value": 187
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": {"$date": 1597794288},
        "value": 4
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "elise_smith@myemail.com",
        "orderdate": {"$date": 1608972946000},
        "value": 4
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": {"$date": 1614570572},
        "value": 1024
    })"),
    bsoncxx::from_json(R"({
        "customer_id": "elise_smith@myemail.com",
        "orderdate": {"$date": 1601722184000},
        "value": 102
    })")
};

auto result = orders.insert_many(docs); // Might throw an exception

原文 C++ 样本和下面的匹配边界混用了秒与毫秒,原文输出也出现了 1970 年日期。使用前应按下文“源码修订说明”统一日期单位,不能把原样输出当作正确的 2020 年统计。

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

pipeline.match(bsoncxx::from_json(R"({
    "orderdate": {
        "$gte": {"$date": 1577836800},
        "$lt": {"$date": 1609459200000}
    }
})"));

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.sort(bsoncxx::from_json(R"({
    "orderdate": 1
})"));

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.group(bsoncxx::from_json(R"({
    "_id": "$customer_id",
    "first_purchase_date": {"$first": "$orderdate"},
    "total_value": {"$sum": "$value"},
    "total_orders": {"$sum": 1},
    "orders": {"$push": {
        "orderdate": "$orderdate",
        "value": "$value"
    }}
})"));

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.sort(bsoncxx::from_json(R"({
    "first_purchase_date": 1
})"));

5. 显示客户邮箱

加入 $addFields 阶段。它是 $set 的别名,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.add_fields(bsoncxx::from_json(R"({
    "customer_id": "$_id"
})"));

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.append_stage(bsoncxx::from_json(R"({
    "$unset": ["_id"]
})"));

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

auto cursor = orders.aggregate(pipeline);

最后在终端中执行以下命令,启动应用:

c++ --std=c++17 agg-tutorial.cpp $(pkg-config --cflags --libs libmongocxx) -o ./app.out
./app.out

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{ "first_purchase_date" : { "$date" : "1970-01-19T06:17:41.137Z" }, "total_value" : 63,
"total_orders" : 1, "orders" : [ { "orderdate" : { "$date" : "1970-01-19T06:17:41.137Z" },
"value" : 63 } ], "customer_id" : "oranieri@warmmail.com" }
{ "first_purchase_date" : { "$date" : "1970-01-19T06:35:01.927Z" }, "total_value" : 436,
"total_orders" : 4, "orders" : [ { "orderdate" : { "$date" : "1970-01-19T06:35:01.927Z" },
"value" : 99 }, { "orderdate" : { "$date" : "2020-05-30T06:55:52Z" }, "value" : 231 },
{ "orderdate" : { "$date" : "2020-10-03T10:49:44Z" }, "value" : 102 }, { "orderdate" :
{ "$date" : "2020-12-26T08:55:46Z" }, "value" : 4 } ], "customer_id" : "elise_smith@myemail.com" }
{ "first_purchase_date" : { "$date" : "1970-01-19T11:49:54.288Z" }, "total_value" : 1215,
"total_orders" : 3, "orders" : [ { "orderdate" : { "$date" : "1970-01-19T11:49:54.288Z" },
"value" : 4 }, { "orderdate" : { "$date" : "1970-01-19T14:09:32.213Z" }, "value" : 187 },
{ "orderdate" : { "$date" : "1970-01-19T16:29:30.572Z" }, "value" : 1024 } ], "customer_id" : "tj@wheresmyemail.com" }

结果按客户邮箱分组,包含这位客户的全部订单详情。

C#

创建模板应用

开始本教程之前,先创建一个新的 C#/.NET 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,把下列代码粘贴到 Program.cs 中,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

using MongoDB.Bson;
using MongoDB.Bson.Serialization.Attributes;
using MongoDB.Driver;

// Define data model classes.
// ... public class MyClass { ... }

// Replace the placeholder with your connection string.
var uri = "<connection string>";
var client = new MongoClient(uri);
var aggDB = client.GetDatabase("agg_tutorials_db");

// Get a reference to relevant collections.
// ... var someColl = aggDB.GetCollection<MyClass>("someColl");
// ... var anotherColl = aggDB.GetCollection<MyClass>("anotherColl");

// Delete any existing documents in collections if needed.
// ... someColl.DeleteMany(Builders<MyClass>.Filter.Empty);
// Insert sample data into the collection or collections.
// ... someColl.InsertMany(new List<MyClass> { ... });

// Add code to chain pipeline stages to the Aggregate() method.
// ... var results = someColl.Aggregate().Match(...);

// Print the aggregation results.
foreach (var result in results.ToList())
{
    Console.WriteLine(result);
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见在 Atlas 中创建免费集群。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

var uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 CustomerId 属性分组。

先创建 C# 类,描述 orders 集合中的数据:

public class Order
{
    [BsonId]
    public ObjectId Id { get; set; }

    public string CustomerId { get; set; } = "";
    public DateTime OrderDate { get; set; }
    public int Value { get; set; }
}

将下列代码加入应用,创建 orders 集合并插入样本数据:

var orders = aggDB.GetCollection<Order>("orders");

orders.InsertMany(new List<Order>
{
    new Order()
    {
        CustomerId = "elise_smith@myemail.com",
        OrderDate = DateTime.Parse("2020-05-30T08:35:52Z"),
        Value = 231
    },
    new Order()
    {
        CustomerId = "elise_smith@myemail.com",
        OrderDate = DateTime.Parse("2020-01-13T09:32:07Z"),
        Value = 99
    },
    new Order()
    {
        CustomerId = "oranieri@warmmail.com",
        OrderDate = DateTime.Parse("2020-01-01T08:25:37Z"),
        Value = 63
    },
    new Order()
    {
        CustomerId = "tj@wheresmyemail.com",
        OrderDate = DateTime.Parse("2019-05-28T19:13:32Z"),
        Value = 2
    },
    new Order()
    {
        CustomerId = "tj@wheresmyemail.com",
        OrderDate = DateTime.Parse("2020-11-23T22:56:53Z"),
        Value = 187
    },
    new Order()
    {
        CustomerId = "tj@wheresmyemail.com",
        OrderDate = DateTime.Parse("2020-08-18T23:04:48Z"),
        Value = 4
    },
    new Order()
    {
        CustomerId = "elise_smith@myemail.com",
        OrderDate = DateTime.Parse("2020-12-26T08:55:46Z"),
        Value = 4
    },
    new Order()
    {
        CustomerId = "tj@wheresmyemail.com",
        OrderDate = DateTime.Parse("2021-02-28T07:49:32Z"),
        Value = 1024
    },
    new Order()
    {
        CustomerId = "elise_smith@myemail.com",
        OrderDate = DateTime.Parse("2020-10-03T13:49:44Z"),
        Value = 102
    }
});

操作步骤

1. 匹配 2020 年的订单

先在 orders 集合上开始聚合,再链接一个 $match 阶段,匹配 2020 年下单的订单。

var results = orders.Aggregate()
    .Match(o => o.OrderDate >= DateTime.Parse("2020-01-01T00:00:00Z") &&
                o.OrderDate < DateTime.Parse("2021-01-01T00:00:00Z"))

2. 按下单日期排序

接着加入 $sort 阶段,按 OrderDate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

.SortBy(o => o.OrderDate)

3. 按邮箱分组

加入 $group 阶段,按 CustomerId 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • CustomerId:客户邮箱,也是分组键。
  • FirstPurchaseDate:客户首次购买的日期。
  • TotalValue:客户所有购买的金额总和。
  • TotalOrders:客户购买的总次数。
  • Orders:客户的全部购买记录,包括每次购买的日期和金额。
.Group(
    id: o => o.CustomerId,
    group: g => new
    {
        CustomerId = g.Key,
        FirstPurchaseDate = g.First().OrderDate,
        TotalValue = g.Sum(i => i.Value),
        TotalOrders = g.Count(),
        Orders = g.Select(i => new { i.OrderDate, i.Value }).ToList()
    }
)

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 FirstPurchaseDate 升序排序。

.SortBy(c => c.FirstPurchaseDate)
.As<BsonDocument>();

上述代码还将输出文档转换为 BsonDocument 实例,便于打印。

5. 运行聚合并解读结果

最后在 IDE 中运行应用并检查结果。原文给出的 2020 年客户订单汇总如下:

{ "CustomerId" : "oranieri@warmmail.com", "FirstPurchaseDate" : { "$date" : "2020-01-01T08:25:37Z" }, "TotalValue" : 63, "TotalOrders" : 1, "Orders" : [{ "OrderDate" : { "$date" : "2020-01-01T08:25:37Z" }, "Value" : 63 }] }
{ "CustomerId" : "elise_smith@myemail.com", "FirstPurchaseDate" : { "$date" : "2020-01-13T09:32:07Z" }, "TotalValue" : 436, "TotalOrders" : 4, "Orders" : [{ "OrderDate" : { "$date" : "2020-01-13T09:32:07Z" }, "Value" : 99 }, { "OrderDate" : { "$date" : "2020-05-30T08:35:52Z" }, "Value" : 231 }, { "OrderDate" : { "$date" : "2020-10-03T13:49:44Z" }, "Value" : 102 }, { "OrderDate" : { "$date" : "2020-12-26T08:55:46Z" }, "Value" : 4 }] }
{ "CustomerId" : "tj@wheresmyemail.com", "FirstPurchaseDate" : { "$date" : "2020-08-18T23:04:48Z" }, "TotalValue" : 191, "TotalOrders" : 2, "Orders" : [{ "OrderDate" : { "$date" : "2020-08-18T23:04:48Z" }, "Value" : 4 }, { "OrderDate" : { "$date" : "2020-11-23T22:56:53Z" }, "Value" : 187 }] }

结果按客户邮箱分组,包含这位客户的全部订单详情。

Go

创建模板应用

开始本教程之前,先创建一个新的 Go 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg_tutorial.go 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

package main

import (
	"context"
	"fmt"
	"log"

	"go.mongodb.org/mongo-driver/v2/bson"
	"go.mongodb.org/mongo-driver/v2/mongo"
	"go.mongodb.org/mongo-driver/v2/mongo/options"
)

// Define structs.
// type MyStruct struct { ... }

func main() {
	ctx := context.Background()
	// Replace the placeholder with your connection string.
	const uri = "<connection string>"

	client, err := mongo.Connect(options.Client().ApplyURI(uri))
	if err != nil {
		log.Fatal(err)
	}

	defer func() {
		if err = client.Disconnect(ctx); err != nil {
			log.Fatal(err)
		}
	}()

	aggDB := client.Database("agg_tutorials_db")

	// Get a reference to relevant collections.
	// ... someColl := aggDB.Collection("...")
	// ... anotherColl := aggDB.Collection("...")

	// Delete any existing documents in collections if needed.
	// ... someColl.DeleteMany(cxt, bson.D{})

	// Insert sample data into the collection or collections.
	// ... _, err = someColl.InsertMany(...)

	// Add code to create pipeline stages.
	// ... myStage := bson.D{{...}}

	// Create a pipeline that includes the stages.
	// ... pipeline := mongo.Pipeline{...}

	// Run the aggregation.
	// ... cursor, err := someColl.Aggregate(ctx, pipeline)

	if err != nil {
		log.Fatal(err)
	}

	defer func() {
		if err := cursor.Close(ctx); err != nil {
			log.Fatalf("failed to close cursor: %v", err)
		}
	}()

	// Decode the aggregation results.
	var results []bson.D
	if err = cursor.All(ctx, &results); err != nil {
		log.Fatalf("failed to decode results: %v", err)
	}

	// Print the aggregation results.
	for _, result := range results {
		res, _ := bson.MarshalExtJSON(result, false, false)
		fmt.Println(string(res))
	}
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建 MongoDB 集群。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

const uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

先创建 Go 结构体,描述 orders 集合中的数据:

type Order struct {
	CustomerID string        `bson:"customer_id,omitempty"`
	OrderDate  bson.DateTime `bson:"orderdate"`
	Value      int           `bson:"value"`
}

将下列代码加入应用,创建 orders 集合并插入样本数据:

orders := aggDB.Collection("orders")
orders.DeleteMany(context.TODO(), bson.D{})

_, err = orders.InsertMany(context.TODO(), []interface{}{
	Order{
		CustomerID: "elise_smith@myemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 5, 30, 8, 35, 53, 0, time.UTC)),
		Value:      231,
	},
	Order{
		CustomerID: "elise_smith@myemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 1, 13, 9, 32, 7, 0, time.UTC)),
		Value:      99,
	},
	Order{
		CustomerID: "oranieri@warmmail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 1, 01, 8, 25, 37, 0, time.UTC)),
		Value:      63,
	},
	Order{
		CustomerID: "tj@wheresmyemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2019, 5, 28, 19, 13, 32, 0, time.UTC)),
		Value:      2,
	},
	Order{
		CustomerID: "tj@wheresmyemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 11, 23, 22, 56, 53, 0, time.UTC)),
		Value:      187,
	},
	Order{
		CustomerID: "tj@wheresmyemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 8, 18, 23, 4, 48, 0, time.UTC)),
		Value:      4,
	},
	Order{
		CustomerID: "elise_smith@myemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 12, 26, 8, 55, 46, 0, time.UTC)),
		Value:      4,
	},
	Order{
		CustomerID: "tj@wheresmyemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2021, 2, 29, 7, 49, 32, 0, time.UTC)),
		Value:      1024,
	},
	Order{
		CustomerID: "elise_smith@myemail.com",
		OrderDate:  bson.NewDateTimeFromTime(time.Date(2020, 10, 3, 13, 49, 44, 0, time.UTC)),
		Value:      102,
	},
})
if err != nil {
	log.Fatal(err)
}

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

matchStage := bson.D{{Key: "$match", Value: bson.D{
	{Key: "orderdate", Value: bson.D{
		{Key: "$gte", Value: time.Date(2020, 1, 1, 0, 0, 0, 0, time.UTC)},
		{Key: "$lt", Value: time.Date(2021, 1, 1, 0, 0, 0, 0, time.UTC)},
	}},
}}}

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

sortStage1 := bson.D{{Key: "$sort", Value: bson.D{
	{Key: "orderdate", Value: 1},
}}}

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
groupStage := bson.D{{Key: "$group", Value: bson.D{
	{Key: "_id", Value: "$customer_id"},
	{Key: "first_purchase_date", Value: bson.D{{Key: "$first", Value: "$orderdate"}}},
	{Key: "total_value", Value: bson.D{{Key: "$sum", Value: "$value"}}},
	{Key: "total_orders", Value: bson.D{{Key: "$sum", Value: 1}}},
	{Key: "orders", Value: bson.D{{Key: "$push", Value: bson.D{
		{Key: "orderdate", Value: "$orderdate"},
		{Key: "value", Value: "$value"},
	}}}},
}}}

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

sortStage2 := bson.D{{Key: "$sort", Value: bson.D{
	{Key: "first_purchase_date", Value: 1},
}}}

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

setStage := bson.D{{Key: "$set", Value: bson.D{
	{Key: "customer_id", Value: "$_id"},
}}}

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

unsetStage := bson.D{{Key: "$unset", Value: bson.A{"_id"}}}

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

pipeline := mongo.Pipeline{matchStage, sortStage1, groupStage, setStage, sortStage2, unsetStage}
cursor, err := orders.Aggregate(context.TODO(), pipeline)

最后在终端中执行以下命令,启动应用:

go run agg_tutorial.go

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{"first_purchase_date":{"$date":"2020-01-01T08:25:37Z"},"total_value":63,"total_orders":1,"orders":[{"orderdate":{"$date":"2020-01-01T08:25:37Z"},"value":63}],"customer_id":"oranieri@warmmail.com"}
{"first_purchase_date":{"$date":"2020-01-13T09:32:07Z"},"total_value":436,"total_orders":4,"orders":[{"orderdate":{"$date":"2020-01-13T09:32:07Z"},"value":99},{"orderdate":{"$date":"2020-05-30T08:35:53Z"},"value":231},{"orderdate":{"$date":"2020-10-03T13:49:44Z"},"value":102},{"orderdate":{"$date":"2020-12-26T08:55:46Z"},"value":4}],"customer_id":"elise_smith@myemail.com"}
{"first_purchase_date":{"$date":"2020-08-18T23:04:48Z"},"total_value":191,"total_orders":2,"orders":[{"orderdate":{"$date":"2020-08-18T23:04:48Z"},"value":4},{"orderdate":{"$date":"2020-11-23T22:56:53Z"},"value":187}],"customer_id":"tj@wheresmyemail.com"}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Java(同步)

创建模板应用

开始本教程之前,先创建一个新的 Java(同步驱动) 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 AggTutorial.java 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

package org.example;

// Modify imports for each tutorial as needed.
import com.mongodb.client.*;
import com.mongodb.client.model.Accumulators;
import com.mongodb.client.model.Aggregates;
import com.mongodb.client.model.Field;
import com.mongodb.client.model.Filters;
import com.mongodb.client.model.Sorts;
import com.mongodb.client.model.Variable;
import org.bson.Document;
import org.bson.conversions.Bson;

import java.time.LocalDateTime;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Collections;
import java.util.List;

public class AggTutorial {
    public static void main(String[] args) {
        // Replace the placeholder with your connection string.
        String uri = "<connection string>";

        try (MongoClient mongoClient = MongoClients.create(uri)) {
            MongoDatabase aggDB = mongoClient.getDatabase("agg_tutorials_db");

            // Get a reference to relevant collections.
            // ... MongoCollection<Document> someColl = ...
            // ... MongoCollection<Document> anotherColl = ...

            // Insert sample data into the collection or collections.
            // ... someColl.insertMany(...);

            // Create an empty pipeline array.
            List<Bson> pipeline = new ArrayList<>();

            // Add code to create pipeline stages.
            // ... pipeline.add(...);

            // Run the aggregation.
            // ... AggregateIterable<Document> aggregationResult =
            // someColl.aggregate(pipeline);

            // Print the aggregation results.
            for (Document document : aggregationResult) {
                System.out.println(document.toJson());
            }
        }
    }
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

String uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

MongoDatabase aggDB = mongoClient.getDatabase("agg_tutorials_db");
MongoCollection<Document> orders = aggDB.getCollection("orders");

orders.insertMany(
        Arrays.asList(
                new Document("customer_id", "elise_smith@myemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-05-30T08:35:52"))
                        .append("value", 231),
                new Document("customer_id", "elise_smith@myemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-01-13T09:32:07"))
                        .append("value", 99),
                new Document("customer_id", "oranieri@warmmail.com")
                        .append("orderdate", LocalDateTime.parse("2020-01-01T08:25:37"))
                        .append("value", 63),
                new Document("customer_id", "tj@wheresmyemail.com")
                        .append("orderdate", LocalDateTime.parse("2019-05-28T19:13:32"))
                        .append("value", 2),
                new Document("customer_id", "tj@wheresmyemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-11-23T22:56:53"))
                        .append("value", 187),
                new Document("customer_id", "tj@wheresmyemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-08-18T23:04:48"))
                        .append("value", 4),
                new Document("customer_id", "elise_smith@myemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-12-26T08:55:46"))
                        .append("value", 4),
                new Document("customer_id", "tj@wheresmyemail.com")
                        .append("orderdate", LocalDateTime.parse("2021-02-28T07:49:32"))
                        .append("value", 1024),
                new Document("customer_id", "elise_smith@myemail.com")
                        .append("orderdate", LocalDateTime.parse("2020-10-03T13:49:44"))
                        .append("value", 102)
        )
);

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

pipeline.add(Aggregates.match(Filters.and(
        Filters.gte("orderdate", LocalDateTime.parse("2020-01-01T00:00:00")),
        Filters.lt("orderdate", LocalDateTime.parse("2021-01-01T00:00:00"))
)));

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.add(Aggregates.sort(Sorts.ascending("orderdate")));

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.add(Aggregates.group(
        "$customer_id",
        Accumulators.first("first_purchase_date", "$orderdate"),
        Accumulators.sum("total_value", "$value"),
        Accumulators.sum("total_orders", 1),
        Accumulators.push("orders",
                new Document("orderdate", "$orderdate")
                        .append("value", "$value")
        )
));

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.add(Aggregates.sort(Sorts.ascending("first_purchase_date")));

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.add(Aggregates.set(new Field<>("customer_id", "$_id")));

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.add(Aggregates.unset("_id"));

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

AggregateIterable<Document> aggregationResult = orders.aggregate(pipeline);

最后在 IDE 中运行应用。

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{"first_purchase_date": {"$date": "2020-01-01T08:25:37Z"}, "total_value": 63, "total_orders": 1, "orders": [{"orderdate": {"$date": "2020-01-01T08:25:37Z"}, "value": 63}], "customer_id": "oranieri@warmmail.com"}
{"first_purchase_date": {"$date": "2020-01-13T09:32:07Z"}, "total_value": 436, "total_orders": 4, "orders": [{"orderdate": {"$date": "2020-01-13T09:32:07Z"}, "value": 99}, {"orderdate": {"$date": "2020-05-30T08:35:52Z"}, "value": 231}, {"orderdate": {"$date": "2020-10-03T13:49:44Z"}, "value": 102}, {"orderdate": {"$date": "2020-12-26T08:55:46Z"}, "value": 4}], "customer_id": "elise_smith@myemail.com"}
{"first_purchase_date": {"$date": "2020-08-18T23:04:48Z"}, "total_value": 191, "total_orders": 2, "orders": [{"orderdate": {"$date": "2020-08-18T23:04:48Z"}, "value": 4}, {"orderdate": {"$date": "2020-11-23T22:56:53Z"}, "value": 187}], "customer_id": "tj@wheresmyemail.com"}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Kotlin(协程)

创建模板应用

开始本教程之前,先创建一个新的 Kotlin(协程驱动) 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

除驱动外,还须把以下依赖加入 build.gradle.kts,然后重新加载项目。

dependencies {
    // Implements Kotlin serialization
    implementation("org.jetbrains.kotlinx:kotlinx-serialization-core:1.5.1")
    // Implements Kotlin date and time handling
    implementation("org.jetbrains.kotlinx:kotlinx-datetime:0.6.1")
}

安装驱动或库后,创建 AggTutorial.kt 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

package org.example

// Modify imports for each tutorial as needed.
import com.mongodb.client.model.*
import com.mongodb.kotlin.client.coroutine.MongoClient
import kotlinx.coroutines.runBlocking
import kotlinx.datetime.LocalDateTime
import kotlinx.datetime.toJavaLocalDateTime
import kotlinx.serialization.Contextual
import kotlinx.serialization.Serializable
import org.bson.Document
import org.bson.conversions.Bson

// Define data classes.
@Serializable
data class MyClass(
    ...
)

suspend fun main() {
    // Replace the placeholder with your connection string.
    val uri = "<connection string>"

    MongoClient.create(uri).use { mongoClient ->
        val aggDB = mongoClient.getDatabase("agg_tutorials_db")

        // Get a reference to relevant collections.
        // ... val someColl = ...

        // Delete any existing documents in collections if needed.
        // ... someColl.deleteMany(empty())

        // Insert sample data into the collection or collections.
        // ... someColl.insertMany( ... )

        // Create an empty pipeline.
        val pipeline = mutableListOf<Bson>()

        // Add code to create pipeline stages.
        // ... pipeline.add(...)

        // Run the aggregation.
        // ... val aggregationResult = someColl.aggregate<Document>(pipeline)

        // Print the aggregation results.
        aggregationResult.collect { println(it) }
    }
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见连接集群。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

val uri = "mongodb+srv://mongodb-example:27017"

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customerID 字段分组。原文叙述在此写作 customer_id,但本分支的数据类及分组代码使用 customerID;customer_id 是后续 $set 阶段创建的输出字段。

先创建 Kotlin 数据类,描述 orders 集合中的数据:

@Serializable
data class Order(
    val customerID: String,
    @Contextual val orderDate: LocalDateTime,
    val value: Int
)

将下列代码加入应用,创建 orders 集合并插入样本数据:

val orders = aggDB.getCollection<Order>("orders")
orders.deleteMany(Filters.empty())

orders.insertMany(
    listOf(
        Order("elise_smith@myemail.com", LocalDateTime.parse("2020-05-30T08:35:52"), 231),
        Order("elise_smith@myemail.com", LocalDateTime.parse("2020-01-13T09:32:07"), 99),
        Order("oranieri@warmmail.com", LocalDateTime.parse("2020-01-01T08:25:37"), 63),
        Order("tj@wheresmyemail.com", LocalDateTime.parse("2019-05-28T19:13:32"), 2),
        Order("tj@wheresmyemail.com", LocalDateTime.parse("2020-11-23T22:56:53"), 187),
        Order("tj@wheresmyemail.com", LocalDateTime.parse("2020-08-18T23:04:48"), 4),
        Order("elise_smith@myemail.com", LocalDateTime.parse("2020-12-26T08:55:46"), 4),
        Order("tj@wheresmyemail.com", LocalDateTime.parse("2021-02-28T07:49:32"), 1024),
        Order("elise_smith@myemail.com", LocalDateTime.parse("2020-10-03T13:49:44"), 102)
    )
)

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

pipeline.add(
    Aggregates.match(
        Filters.and(
            Filters.gte(Order::orderDate.name, LocalDateTime.parse("2020-01-01T00:00:00").toJavaLocalDateTime()),
            Filters.lt(Order::orderDate.name, LocalDateTime.parse("2021-01-01T00:00:00").toJavaLocalDateTime())
        )
    )
)

2. 按下单日期排序

接着加入 $sort 阶段,按 orderDate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.add(Aggregates.sort(Sorts.ascending(Order::orderDate.name)))

3. 按邮箱分组

加入 $group 阶段,按 customerID 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.add(
    Aggregates.group(
        "\$${Order::customerID.name}",
        Accumulators.first("first_purchase_date", "\$${Order::orderDate.name}"),
        Accumulators.sum("total_value", "\$${Order::value.name}"),
        Accumulators.sum("total_orders", 1),
        Accumulators.push(
            "orders",
            Document("orderdate", "\$${Order::orderDate.name}")
                .append("value", "\$${Order::value.name}")
        )
    )
)

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.add(Aggregates.sort(Sorts.ascending("first_purchase_date")))

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.add(Aggregates.set(Field("customer_id", "\$_id")))

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.add(Aggregates.unset("_id"))

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

val aggregationResult = orders.aggregate<Document>(pipeline)

最后在 IDE 中运行应用。

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

Document{{first_purchase_date=Wed Jan 01 03:25:37 EST 2020, total_value=63, total_orders=1, orders=[Document{{orderdate=Wed Jan 01 03:25:37 EST 2020, value=63}}], customer_id=oranieri@warmmail.com}}
Document{{first_purchase_date=Mon Jan 13 04:32:07 EST 2020, total_value=436, total_orders=4, orders=[Document{{orderdate=Mon Jan 13 04:32:07 EST 2020, value=99}}, Document{{orderdate=Sat May 30 04:35:52 EDT 2020, value=231}}, Document{{orderdate=Sat Oct 03 09:49:44 EDT 2020, value=102}}, Document{{orderdate=Sat Dec 26 03:55:46 EST 2020, value=4}}], customer_id=elise_smith@myemail.com}}
Document{{first_purchase_date=Tue Aug 18 19:04:48 EDT 2020, total_value=191, total_orders=2, orders=[Document{{orderdate=Tue Aug 18 19:04:48 EDT 2020, value=4}}, Document{{orderdate=Mon Nov 23 17:56:53 EST 2020, value=187}}], customer_id=tj@wheresmyemail.com}}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Node.js

创建模板应用

开始本教程之前,先创建一个新的 Node.js 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动后,创建用于运行教程模板的文件,并把以下代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

import { MongoClient } from 'mongodb';

// Replace the placeholder with your connection string.
const uri = '<connection-string>';
const client = new MongoClient(uri);
export async function run() {
  try {
    const aggDB = client.db('agg_tutorials_db');

    // Get a reference to relevant collections.
    // ... const someColl =
    // ... const anotherColl =

    // Delete any existing documents in collections.
    // ... await someColl.deleteMany({});

    // Insert sample data into the collection or collections.
    // ... const someData = [ ... ];
    // ... await someColl.insertMany(someData);

    // Create an empty pipeline array.
    const pipeline = [];

    // Add code to create pipeline stages.
    // ... pipeline.push({ ... })

    // Run the aggregation.
    // ... const aggregationResult = ...

    // Print the aggregation results.
    for await (const document of aggregationResult) {
      console.log(document);
    }
  } finally {
    await client.close();
  }
}

run().catch(console.dir);

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

const uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

const orders = aggDB.collection('orders');

await orders.insertMany([
  {
    customer_id: 'elise_smith@myemail.com',
    orderdate: new Date('2020-05-30T08:35:52Z'),
    value: 231,
  },
  {
    customer_id: 'elise_smith@myemail.com',
    orderdate: new Date('2020-01-13T09:32:07Z'),
    value: 99,
  },
  {
    customer_id: 'oranieri@warmmail.com',
    orderdate: new Date('2020-01-01T08:25:37Z'),
    value: 63,
  },
  {
    customer_id: 'tj@wheresmyemail.com',
    orderdate: new Date('2019-05-28T19:13:32Z'),
    value: 2,
  },
  {
    customer_id: 'tj@wheresmyemail.com',
    orderdate: new Date('2020-11-23T22:56:53Z'),
    value: 187,
  },
  {
    customer_id: 'tj@wheresmyemail.com',
    orderdate: new Date('2020-08-18T23:04:48Z'),
    value: 4,
  },
  {
    customer_id: 'elise_smith@myemail.com',
    orderdate: new Date('2020-12-26T08:55:46Z'),
    value: 4,
  },
  {
    customer_id: 'tj@wheresmyemail.com',
    orderdate: new Date('2021-02-29T07:49:32Z'),
    value: 1024,
  },
  {
    customer_id: 'elise_smith@myemail.com',
    orderdate: new Date('2020-10-03T13:49:44Z'),
    value: 102,
  },
]);

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

pipeline.push({
  $match: {
    orderdate: {
      $gte: new Date('2020-01-01T00:00:00Z'),
      $lt: new Date('2021-01-01T00:00:00Z'),
    },
  },
});

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.push({
  $sort: {
    orderdate: 1,
  },
});

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.push({
  $group: {
    _id: '$customer_id',
    first_purchase_date: { $first: '$orderdate' },
    total_value: { $sum: '$value' },
    total_orders: { $sum: 1 },
    orders: {
      $push: {
        orderdate: '$orderdate',
        value: '$value',
      },
    },
  },
});

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.push({
  $sort: {
    first_purchase_date: 1,
  },
});

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.push({
  $set: {
    customer_id: '$_id',
  },
});

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.push({ $unset: ['_id'] });

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

const aggregationResult = await orders.aggregate(pipeline);

最后使用 IDE 或命令行执行该文件中的代码。

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{
  first_purchase_date: 2020-01-01T08:25:37.000Z,
  total_value: 63,
  total_orders: 1,
  orders: [ { orderdate: 2020-01-01T08:25:37.000Z, value: 63 } ],
  customer_id: 'oranieri@warmmail.com'
}
{
  first_purchase_date: 2020-01-13T09:32:07.000Z,
  total_value: 436,
  total_orders: 4,
  orders: [
    { orderdate: 2020-01-13T09:32:07.000Z, value: 99 },
    { orderdate: 2020-05-30T08:35:52.000Z, value: 231 },
    { orderdate: 2020-10-03T13:49:44.000Z, value: 102 },
    { orderdate: 2020-12-26T08:55:46.000Z, value: 4 }
  ],
  customer_id: 'elise_smith@myemail.com'
}
{
  first_purchase_date: 2020-08-18T23:04:48.000Z,
  total_value: 191,
  total_orders: 2,
  orders: [
    { orderdate: 2020-08-18T23:04:48.000Z, value: 4 },
    { orderdate: 2020-11-23T22:56:53.000Z, value: 187 }
  ],
  customer_id: 'tj@wheresmyemail.com'
}

结果按客户邮箱分组,包含这位客户的全部订单详情。

PHP

创建模板应用

开始本教程之前,先创建一个新的 PHP 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg_tutorial.php 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

<?php

require 'vendor/autoload.php';

// Modify imports for each tutorial as needed.
use MongoDB\Client;
use MongoDB\BSON\UTCDateTime;
use MongoDB\Builder\Pipeline;
use MongoDB\Builder\Stage;
use MongoDB\Builder\Type\Sort;
use MongoDB\Builder\Query;
use MongoDB\Builder\Expression;
use MongoDB\Builder\Accumulator;

use function MongoDB\object;

// Replace the placeholder with your connection string.
$uri = '<connection string>';
$client = new Client($uri);

// Get a reference to relevant collections.
// ... $someColl = $client->agg_tutorials_db->someColl;
// ... $anotherColl = $client->agg_tutorials_db->anotherColl;

// Delete any existing documents in collections if needed.
// ... $someColl->deleteMany([]);

// Insert sample data into the collection or collections.
// ... $someColl->insertMany(...);

// Add code to create pipeline stages within the Pipeline instance.
// ... $pipeline = new Pipeline(...);

// Run the aggregation.
// ... $cursor = $someColl->aggregate($pipeline);

// Print the aggregation results.
foreach ($cursor as $doc) {
    echo json_encode($doc, JSON_PRETTY_PRINT), PHP_EOL;
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

$uri = 'mongodb+srv://mongodb-example:27017';

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

$orders = $client->agg_tutorials_db->orders;
$orders->deleteMany([]);

$orders->insertMany(
    [
        [
            'customer_id' => 'elise_smith@myemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-05-30T08:35:52')),
            'value' => 231
        ],
        [
            'customer_id' => 'elise_smith@myemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-01-13T09:32:07')),
            'value' => 99
        ],
        [
            'customer_id' => 'oranieri@warmmail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-01-01T08:25:37')),
            'value' => 63
        ],
        [
            'customer_id' => 'tj@wheresmyemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2019-05-28T19:13:32')),
            'value' => 2
        ],
        [
            'customer_id' => 'tj@wheresmyemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-11-23T22:56:53')),
            'value' => 187
        ],
        [
            'customer_id' => 'tj@wheresmyemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-08-18T23:04:48')),
            'value' => 4
        ],
        [
            'customer_id' => 'elise_smith@myemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-12-26T08:55:46')),
            'value' => 4
        ],
        [
            'customer_id' => 'tj@wheresmyemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2021-02-28T07:49:32')),
            'value' => 1024
        ],
        [
            'customer_id' => 'elise_smith@myemail.com',
            'orderdate' => new UTCDateTime(new DateTimeImmutable('2020-10-03T13:49:44')),
            'value' => 102
        ]
    ]
);

操作步骤

1. 匹配 2020 年的订单

在 Pipeline 实例中加入 $match 阶段,匹配 2020 年下单的订单。

Stage::match(
    orderdate: [
        Query::gte(new UTCDateTime(new DateTimeImmutable('2020-01-01T00:00:00'))),
        Query::lt(new UTCDateTime(new DateTimeImmutable('2021-01-01T00:00:00'))),
    ]
),

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

Stage::sort(orderdate: Sort::Asc),

3. 按邮箱分组

在 Pipeline 实例外创建工厂函数,返回 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
function groupByCustomerStage()
{
    return Stage::group(
        _id: Expression::stringFieldPath('customer_id'),
        first_purchase_date: Accumulator::first(
            Expression::arrayFieldPath('orderdate')
        ),
        total_value: Accumulator::sum(
            Expression::numberFieldPath('value'),
        ),
        total_orders: Accumulator::sum(1),
        orders: Accumulator::push(
            object(
                orderdate: Expression::dateFieldPath('orderdate'),
                value: Expression::numberFieldPath('value'),
            ),
        ),
    );
}

然后在 Pipeline 实例中调用 groupByCustomerStage():

groupByCustomerStage(),

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

Stage::sort(first_purchase_date: Sort::Asc),

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

Stage::set(customer_id: Expression::stringFieldPath('_id')),

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

Stage::unset('_id')

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

$cursor = $orders->aggregate($pipeline);

最后在终端中执行以下命令,启动应用:

php agg_tutorial.php

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{
    "first_purchase_date": {
        "$date": {
            "$numberLong": "1577867137000"
        }
    },
    "total_value": 63,
    "total_orders": 1,
    "orders": [
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1577867137000"
                }
            },
            "value": 63
        }
    ],
    "customer_id": "oranieri@warmmail.com"
}
{
    "first_purchase_date": {
        "$date": {
            "$numberLong": "1578907927000"
        }
    },
    "total_value": 436,
    "total_orders": 4,
    "orders": [
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1578907927000"
                }
            },
            "value": 99
        },
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1590827752000"
                }
            },
            "value": 231
        },
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1601732984000"
                }
            },
            "value": 102
        },
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1608972946000"
                }
            },
            "value": 4
        }
    ],
    "customer_id": "elise_smith@myemail.com"
}
{
    "first_purchase_date": {
        "$date": {
            "$numberLong": "1597791888000"
        }
    },
    "total_value": 191,
    "total_orders": 2,
    "orders": [
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1597791888000"
                }
            },
            "value": 4
        },
        {
            "orderdate": {
                "$date": {
                    "$numberLong": "1606172213000"
                }
            },
            "value": 187
        }
    ],
    "customer_id": "tj@wheresmyemail.com"
}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Python

创建模板应用

开始本教程之前,先创建一个新的 Python / PyMongo 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg_tutorial.py 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

# Modify imports for each tutorial as needed.
from pymongo import MongoClient

# Replace the placeholder with your connection string.
uri = "<connection-string>"
client = MongoClient(uri)

try:
    agg_db = client["agg_tutorials_db"]

    # Get a reference to relevant collections.
    # ... some_coll = agg_db["some_coll"]
    # ... another_coll = agg_db["another_coll"]

    # Delete any existing documents in collections if needed.
    # ... some_coll.delete_many({})

    # Insert sample data into the collection or collections.
    # ... some_coll.insert_many(...)

    # Create an empty pipeline array.
    pipeline = []

    # Add code to create pipeline stages.
    # ... pipeline.append({...})

    # Run the aggregation.
    # ... aggregation_result = ...

    # Print the aggregation results.
    for document in aggregation_result:
        print(document)

finally:
    client.close()

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见原文引用的连接字符串指引(PHP)。

原文在此引用了 PHP 库的连接指引;PyMongo 的安装与连接说明请同时参照上面的 Python 入门指南。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

uri = "mongodb+srv://mongodb-example:27017"

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

orders_coll = agg_db["orders"]

order_data = [
    {
        "customer_id": "elise_smith@myemail.com",
        "orderdate": datetime(2020, 5, 30, 8, 35, 52),
        "value": 231,
    },
    {
        "customer_id": "elise_smith@myemail.com",
        "orderdate": datetime(2020, 1, 13, 9, 32, 7),
        "value": 99,
    },
    {
        "customer_id": "oranieri@warmmail.com",
        "orderdate": datetime(2020, 1, 1, 8, 25, 37),
        "value": 63,
    },
    {
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": datetime(2019, 5, 28, 19, 13, 32),
        "value": 2,
    },
    {
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": datetime(2020, 11, 23, 22, 56, 53),
        "value": 187,
    },
    {
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": datetime(2020, 8, 18, 23, 4, 48),
        "value": 4,
    },
    {
        "customer_id": "elise_smith@myemail.com",
        "orderdate": datetime(2020, 12, 26, 8, 55, 46),
        "value": 4,
    },
    {
        "customer_id": "tj@wheresmyemail.com",
        "orderdate": datetime(2021, 2, 28, 7, 49, 32),
        "value": 1024,
    },
    {
        "customer_id": "elise_smith@myemail.com",
        "orderdate": datetime(2020, 10, 3, 13, 49, 44),
        "value": 102,
    },
]

orders_coll.insert_many(order_data)

操作步骤

1. 匹配 2020 年的订单

原文要求在管道中加入 $match 阶段,匹配 2020 年下单的订单;该 Python 示例通过 pipeline.append(...) 构建阶段列表。

pipeline.append(
    {
        "$match": {
            "orderdate": {
                "$gte": datetime(2020, 1, 1, 0, 0, 0),
                "$lt": datetime(2021, 1, 1, 0, 0, 0),
            }
        }
    }
)

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.append({"$sort": {"orderdate": 1}})

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.append(
    {
        "$group": {
            "_id": "$customer_id",
            "first_purchase_date": {"$first": "$orderdate"},
            "total_value": {"$sum": "$value"},
            "total_orders": {"$sum": 1},
            "orders": {"$push": {"orderdate": "$orderdate", "value": "$value"}},
        }
    }
)

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.append({"$sort": {"first_purchase_date": 1}})

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.append({"$set": {"customer_id": "$_id"}})

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.append({"$unset": ["_id"]})

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

aggregation_result = orders_coll.aggregate(pipeline)

最后在终端中执行以下命令,启动应用:

python3 agg_tutorial.py

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{'first_purchase_date': datetime.datetime(2020, 1, 1, 8, 25, 37), 'total_value': 63, 'total_orders': 1, 'orders': [{'orderdate': datetime.datetime(2020, 1, 1, 8, 25, 37), 'value': 63}], 'customer_id': 'oranieri@warmmail.com'}
{'first_purchase_date': datetime.datetime(2020, 1, 13, 9, 32, 7), 'total_value': 436, 'total_orders': 4, 'orders': [{'orderdate': datetime.datetime(2020, 1, 13, 9, 32, 7), 'value': 99}, {'orderdate': datetime.datetime(2020, 5, 30, 8, 35, 52), 'value': 231}, {'orderdate': datetime.datetime(2020, 10, 3, 13, 49, 44), 'value': 102}, {'orderdate': datetime.datetime(2020, 12, 26, 8, 55, 46), 'value': 4}], 'customer_id': 'elise_smith@myemail.com'}
{'first_purchase_date': datetime.datetime(2020, 8, 18, 23, 4, 48), 'total_value': 191, 'total_orders': 2, 'orders': [{'orderdate': datetime.datetime(2020, 8, 18, 23, 4, 48), 'value': 4}, {'orderdate': datetime.datetime(2020, 11, 23, 22, 56, 53), 'value': 187}], 'customer_id': 'tj@wheresmyemail.com'}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Ruby

创建模板应用

开始本教程之前,先创建一个新的 Ruby 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg_tutorial.rb 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

# typed: strict
require 'mongo'
require 'bson'

# Replace the placeholder with your connection string.
uri = "<connection string>"

Mongo::Client.new(uri) do |client|

  agg_db = client.use('agg_tutorials_db')

  # Get a reference to relevant collections.
  # ... some_coll = agg_db[:some_coll]

  # Delete any existing documents in collections if needed.
  # ... some_coll.delete_many({})

  # Insert sample data into the collection or collections.
  # ... some_coll.insert_many( ... )

  # Add code to create pipeline stages within the array.
  # ... pipeline = [ ... ]

  # Run the aggregation.
  # ... aggregation_result = some_coll.aggregate(pipeline)

  # Print the aggregation results.
  aggregation_result.each do |doc|
    puts doc
  end

end

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

uri = "mongodb+srv://mongodb-example:27017"

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

orders = agg_db[:orders]
orders.delete_many({})

orders.insert_many(
  [
    {
      customer_id: "elise_smith@myemail.com",
      orderdate: DateTime.parse("2020-05-30T08:35:52Z"),
      value: 231,

    },
    {
      customer_id: "elise_smith@myemail.com",
      orderdate: DateTime.parse("2020-01-13T09:32:07Z"),
      value: 99,
    },
    {
      customer_id: "oranieri@warmmail.com",
      orderdate: DateTime.parse("2020-01-01T08:25:37Z"),
      value: 63,
    },
    {
      customer_id: "tj@wheresmyemail.com",
      orderdate: DateTime.parse("2019-05-28T19:13:32Z"),
      value: 2,
    },
    {
      customer_id: "tj@wheresmyemail.com",
      orderdate: DateTime.parse("2020-11-23T22:56:53Z"),
      value: 187,
    },
    {
      customer_id: "tj@wheresmyemail.com",
      orderdate: DateTime.parse("2020-08-18T23:04:48Z"),
      value: 4,
    },
    {
      customer_id: "elise_smith@myemail.com",
      orderdate: DateTime.parse("2020-12-26T08:55:46Z"),
      value: 4,
    },
    {
      customer_id: "tj@wheresmyemail.com",
      orderdate: DateTime.parse("2021-02-28T07:49:32Z"),
      value: 1024,
    },
    {
      customer_id: "elise_smith@myemail.com",
      orderdate: DateTime.parse("2020-10-03T13:49:44Z"),
      value: 102,
    },
  ]
)

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

{
  "$match": {
    orderdate: {
      "$gte": DateTime.parse("2020-01-01T00:00:00Z"),
      "$lt": DateTime.parse("2021-01-01T00:00:00Z"),
    },
  },
},

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

{
  "$sort": {
    orderdate: 1,
  },
},

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
{
  "$group": {
    _id: "$customer_id",
    first_purchase_date: { "$first": "$orderdate" },
    total_value: { "$sum": "$value" },
    total_orders: { "$sum": 1 },
    orders: { "$push": {
      orderdate: "$orderdate",
      value: "$value",
    } },
  },
},

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

{
  "$sort": {
    first_purchase_date: 1,
  },
},

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

{
  "$set": {
    customer_id: "$_id",
  },
},

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

{ "$unset": ["_id"] },

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

aggregation_result = orders.aggregate(pipeline)

最后在终端中执行以下命令,启动应用:

ruby agg_tutorial.rb

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{"first_purchase_date"=>2020-01-01 08:25:37 UTC, "total_value"=>63, "total_orders"=>1, "orders"=>[{"orderdate"=>2020-01-01 08:25:37 UTC, "value"=>63}], "customer_id"=>"oranieri@warmmail.com"}
{"first_purchase_date"=>2020-01-13 09:32:07 UTC, "total_value"=>436, "total_orders"=>4, "orders"=>[{"orderdate"=>2020-01-13 09:32:07 UTC, "value"=>99}, {"orderdate"=>2020-05-30 08:35:52 UTC, "value"=>231}, {"orderdate"=>2020-10-03 13:49:44 UTC, "value"=>102}, {"orderdate"=>2020-12-26 08:55:46 UTC, "value"=>4}], "customer_id"=>"elise_smith@myemail.com"}
{"first_purchase_date"=>2020-08-18 23:04:48 UTC, "total_value"=>191, "total_orders"=>2, "orders"=>[{"orderdate"=>2020-08-18 23:04:48 UTC, "value"=>4}, {"orderdate"=>2020-11-23 22:56:53 UTC, "value"=>187}], "customer_id"=>"tj@wheresmyemail.com"}

结果按客户邮箱分组,包含这位客户的全部订单详情。

Rust

创建模板应用

开始本教程之前,先创建一个新的 Rust 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 agg-tutorial.rs 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

use mongodb::{
    bson::{doc, Document},
    options::ClientOptions,
    Client,
};
use futures::stream::TryStreamExt;
use std::error::Error;

// Define structs.
// #[derive(Debug, Serialize, Deserialize)]
// struct MyStruct { ... }

#[tokio::main]
async fn main() mongodb::error::Result<()> {
    // Replace the placeholder with your connection string.
    let uri = "<connection string>";
    let client = Client::with_uri_str(uri).await?;

    let agg_db = client.database("agg_tutorials_db");

    // Get a reference to relevant collections.
    // ... let some_coll: Collection<T> = agg_db.collection("...");
    // ... let another_coll: Collection<T> = agg_db.collection("...");

    // Delete any existing documents in collections if needed.
    // ... some_coll.delete_many(doc! {}).await?;

    // Insert sample data into the collection or collections.
    // ... some_coll.insert_many(vec![...]).await?;

    // Create an empty pipeline.
    let mut pipeline = Vec::new();

    // Add code to create pipeline stages.
    // pipeline.push(doc! { ... });

    // Run the aggregation and print the results.
    let mut results = some_coll.aggregate(pipeline).await?;
    while let Some(result) = results.try_next().await? {
       println!("{:?}\n", result);
    }
    Ok(())
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

let uri = "mongodb+srv://mongodb-example:27017";

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

先创建 Rust 结构体,描述 orders 集合中的数据:

#[derive(Debug, Serialize, Deserialize)]
struct Order {
    customer_id: String,
    orderdate: DateTime,
    value: i32,
}

将下列代码加入应用,创建 orders 集合并插入样本数据:

let orders: Collection<Order> = agg_db.collection("orders");
orders.delete_many(doc! {}).await?;

let docs = vec![
    Order {
        customer_id: "elise_smith@myemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(5).day(30).hour(8).minute(35).second(53).build().unwrap(),
        value: 231,
    },
    Order {
        customer_id: "elise_smith@myemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(1).day(13).hour(9).minute(32).second(7).build().unwrap(),
        value: 99,
    },
    Order {
        customer_id: "oranieri@warmmail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(1).day(1).hour(8).minute(25).second(37).build().unwrap(),
        value: 63,
    },
    Order {
        customer_id: "tj@wheresmyemail.com".to_string(),
        orderdate: DateTime::builder().year(2019).month(5).day(28).hour(19).minute(13).second(32).build().unwrap(),
        value: 2,
    },
    Order {
        customer_id: "tj@wheresmyemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(11).day(23).hour(22).minute(56).second(53).build().unwrap(),
        value: 187,
    },
    Order {
        customer_id: "tj@wheresmyemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(8).day(18).hour(23).minute(4).second(48).build().unwrap(),
        value: 4,
    },
    Order {
        customer_id: "elise_smith@myemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(12).day(26).hour(8).minute(55).second(46).build().unwrap(),
        value: 4,
    },
    Order {
        customer_id: "tj@wheresmyemail.com".to_string(),
        orderdate: DateTime::builder().year(2021).month(2).day(28).hour(7).minute(49).second(32).build().unwrap(),
        value: 1024,
    },
    Order {
        customer_id: "elise_smith@myemail.com".to_string(),
        orderdate: DateTime::builder().year(2020).month(10).day(3).hour(13).minute(49).second(44).build().unwrap(),
        value: 102,
    },
];

orders.insert_many(docs).await?;

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

pipeline.push(doc! {
    "$match": {
        "orderdate": {
            "$gte": DateTime::builder().year(2020).month(1).day(1).build().unwrap(),
            "$lt": DateTime::builder().year(2021).month(1).day(1).build().unwrap(),
        }
    }
});

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

pipeline.push(doc! {
    "$sort": {
        "orderdate": 1
    }
});

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
pipeline.push(doc! {
    "$group": {
        "_id": "$customer_id",
        "first_purchase_date": { "$first": "$orderdate" },
        "total_value": { "$sum": "$value" },
        "total_orders": { "$sum": 1 },
        "orders": {
            "$push": {
                "orderdate": "$orderdate",
                "value": "$value"
            }
        }
    }
});

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

pipeline.push(doc! {
    "$sort": {
        "first_purchase_date": 1
    }
});

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

pipeline.push(doc! {
    "$set": {
        "customer_id": "$_id"
    }
});

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

pipeline.push(doc! {"$unset": ["_id"] });

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

let mut cursor = orders.aggregate(pipeline).await?;

最后在终端中执行以下命令,启动应用:

cargo run

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

Document({"first_purchase_date": DateTime(2020-01-01 8:25:37.0 +00:00:00), "total_value": Int32(63), "total_orders": Int32(1),
"orders": Array([Document({"orderdate": DateTime(2020-01-01 8:25:37.0 +00:00:00), "value": Int32(63)})]), "customer_id": String("oranieri@warmmail.com")})

Document({"first_purchase_date": DateTime(2020-01-13 9:32:07.0 +00:00:00), "total_value": Int32(436), "total_orders": Int32(4),
"orders": Array([Document({"orderdate": DateTime(2020-01-13 9:32:07.0 +00:00:00), "value": Int32(99)}), Document({"orderdate":
DateTime(2020-05-30 8:35:53.0 +00:00:00), "value": Int32(231)}), Document({"orderdate": DateTime(2020-10-03 13:49:44.0 +00:00:00),
"value": Int32(102)}), Document({"orderdate": DateTime(2020-12-26 8:55:46.0 +00:00:00), "value": Int32(4)})]), "customer_id": String("elise_smith@myemail.com")})

Document({"first_purchase_date": DateTime(2020-08-18 23:04:48.0 +00:00:00), "total_value": Int32(191), "total_orders": Int32(2),
"orders": Array([Document({"orderdate": DateTime(2020-08-18 23:04:48.0 +00:00:00), "value": Int32(4)}), Document({"orderdate":
DateTime(2020-11-23 22:56:53.0 +00:00:00), "value": Int32(187)})]), "customer_id": String("tj@wheresmyemail.com")})

结果按客户邮箱分组,包含这位客户的全部订单详情。

Scala

创建模板应用

开始本教程之前,先创建一个新的 Scala 应用。它用于连接 MongoDB 部署、插入样本数据,并执行聚合管道。

驱动或库的安装和连接方法见入门指南;聚合方法见聚合指南。

安装驱动或库后,创建 AggTutorial.scala 文件,并把下列代码粘贴进去,建立聚合教程的应用模板。

阅读代码注释,找到需要根据本教程补充或修改的位置。如果不修改模板便尝试运行,会遇到连接错误。模板中的占位集合、导入和管道也需要按对应示例补齐。

package org.example;

// Modify imports for each tutorial as needed.
import org.mongodb.scala.MongoClient
import org.mongodb.scala.bson.Document
import org.mongodb.scala.model.{Accumulators, Aggregates, Field, Filters, Variable}

import java.text.SimpleDateFormat

object FilteredSubset {

  def main(args: Array[String]): Unit = {

    // Replace the placeholder with your connection string.
    val uri = "<connection string>"
    val mongoClient = MongoClient(uri)
    Thread.sleep(1000)

    val aggDB = mongoClient.getDatabase("agg_tutorials_db")

    // Get a reference to relevant collections.
    // ... val someColl = aggDB.getCollection("someColl")
    // ... val anotherColl = aggDB.getCollection("anotherColl")

    // Delete any existing documents in collections if needed.
    // ... someColl.deleteMany(Filters.empty()).subscribe(...)

    // If needed, create the date format template.
    val dateFormat = new SimpleDateFormat("yyyy-MM-dd'T'HH:mm:ss")

    // Insert sample data into the collection or collections.
    // ... someColl.insertMany(...).subscribe(...)

    Thread.sleep(1000)

    // Add code to create pipeline stages within the Seq.
    // ... val pipeline = Seq(...)

    // Run the aggregation and print the results.
    // ... someColl.aggregate(pipeline).subscribe(...)

    Thread.sleep(1000)
    mongoClient.close()
  }
}

每个教程都必须把连接字符串占位符替换为你自己的部署连接字符串。获取方法见创建连接字符串。

例如,若连接字符串为 "mongodb+srv://mongodb-example:27017",赋值语句如下。这个字符串只是原文的示例占位值。

val uri = "mongodb+srv://mongodb-example:27017"

创建集合

本例的 orders 集合保存单个商品订单。每个订单对应一位客户,因此按包含客户邮箱的 customer_id 字段分组。

将下列代码加入应用,创建 orders 集合并插入样本数据:

val orders = aggDB.getCollection("orders")

orders.deleteMany(Filters.empty()).subscribe(
  _ => {},
  e => println("Error: " + e.getMessage),
)

val dateFormat = new SimpleDateFormat("yyyy-MM-dd'T'HH:mm:ss")

orders.insertMany(Seq(
  Document("customer_id" -> "elise_smith@myemail.com",
    "orderdate" -> dateFormat.parse("2020-05-30T08:35:52"),
    "value" -> 231),
  Document("customer_id" -> "elise_smith@myemail.com",
    "orderdate" -> dateFormat.parse("2020-01-13T09:32:07"),
    "value" -> 99),
  Document("customer_id" -> "oranieri@warmmail.com",
    "orderdate" -> dateFormat.parse("2020-01-01T08:25:37"),
    "value" -> 63),
  Document("customer_id" -> "tj@wheresmyemail.com",
    "orderdate" -> dateFormat.parse("2019-05-28T19:13:32"),
    "value" -> 2),
  Document("customer_id" -> "tj@wheresmyemail.com",
    "orderdate" -> dateFormat.parse("2020-11-23T22:56:53"),
    "value" -> 187),
  Document("customer_id" -> "tj@wheresmyemail.com",
    "orderdate" -> dateFormat.parse("2020-08-18T23:04:48"),
    "value" -> 4),
  Document("customer_id" -> "elise_smith@myemail.com",
    "orderdate" -> dateFormat.parse("2020-12-26T08:55:46"),
    "value" -> 4),
  Document("customer_id" -> "tj@wheresmyemail.com",
    "orderdate" -> dateFormat.parse("2021-02-28T07:49:32"),
    "value" -> 1024),
  Document("customer_id" -> "elise_smith@myemail.com",
    "orderdate" -> dateFormat.parse("2020-10-03T13:49:44"),
    "value" -> 102)
)).subscribe(
  _ => {},
  e => println("Error: " + e.getMessage),
)

操作步骤

1. 匹配 2020 年的订单

先加入 $match 阶段,匹配 2020 年下单的订单。

Aggregates.filter(Filters.and(
  Filters.gte("orderdate", dateFormat.parse("2020-01-01T00:00:00")),
  Filters.lt("orderdate", dateFormat.parse("2021-01-01T00:00:00"))
)),

2. 按下单日期排序

接着加入 $sort 阶段,按 orderdate 升序排序,让下一个阶段能够取得每位客户在 2020 年最早的一次购买。

Aggregates.sort(Sorts.ascending("orderdate")),

3. 按邮箱分组

加入 $group 阶段,按 customer_id 的值收集订单文档。加入聚合运算,在结果文档中建立以下字段:

  • first_purchase_date:客户首次购买的日期。
  • total_value:客户所有购买的金额总和。
  • total_orders:客户购买的总次数。
  • orders:客户的全部购买记录,包括每次购买的日期和金额。
Aggregates.group(
  "$customer_id",
  Accumulators.first("first_purchase_date", "$orderdate"),
  Accumulators.sum("total_value", "$value"),
  Accumulators.sum("total_orders", 1),
  Accumulators.push("orders", Document("orderdate" -> "$orderdate", "value" -> "$value"))
),

4. 按首次购买日期排序

再加入一个 $sort 阶段,按 first_purchase_date 升序排序。

Aggregates.sort(Sorts.ascending("first_purchase_date")),

5. 显示客户邮箱

加入 $set 阶段,使用 $group 阶段写入 _id 的值重新建立 customer_id 字段。

Aggregates.set(Field("customer_id", "$_id")),

6. 移除不需要的字段

最后加入 $unset 阶段,从结果文档中移除 _id 字段。

Aggregates.unset("_id")

7. 运行聚合管道

把下列代码加入应用末尾,对 orders 集合执行聚合。

orders.aggregate(pipeline)
  .subscribe((doc: Document) => println(doc.toJson()),
    (e: Throwable) => println(s"Error: $e"))

最后在 IDE 中运行应用。

8. 解读聚合结果

原文给出的 2020 年客户订单汇总如下:

{"first_purchase_date": {"$date": "2020-01-01T13:25:37Z"}, "total_value": 63, "total_orders": 1, "orders": [{"orderdate": {"$date": "2020-01-01T13:25:37Z"}, "value": 63}], "customer_id": "oranieri@warmmail.com"}
{"first_purchase_date": {"$date": "2020-01-13T14:32:07Z"}, "total_value": 436, "total_orders": 4, "orders": [{"orderdate": {"$date": "2020-01-13T14:32:07Z"}, "value": 99}, {"orderdate": {"$date": "2020-05-30T12:35:52Z"}, "value": 231}, {"orderdate": {"$date": "2020-10-03T17:49:44Z"}, "value": 102}, {"orderdate": {"$date": "2020-12-26T13:55:46Z"}, "value": 4}], "customer_id": "elise_smith@myemail.com"}
{"first_purchase_date": {"$date": "2020-08-19T03:04:48Z"}, "total_value": 191, "total_orders": 2, "orders": [{"orderdate": {"$date": "2020-08-19T03:04:48Z"}, "value": 4}, {"orderdate": {"$date": "2020-11-24T03:56:53Z"}, "value": 187}], "customer_id": "tj@wheresmyemail.com"}

结果按客户邮箱分组,包含这位客户的全部订单详情。

源码修订说明

上述代码与输出保留原文内容。下面列出应在实际使用前处理的问题;这些说明不代表已运行各驱动示例。

  • MongoDB Shell 和 Node.js 样本中的 2021-02-29 不是有效日期。若意图使用二月底,可将其修为 2021-02-28;无论采用哪个有效的 2021 年日期,都应落在本例的 2020 年筛选范围之外。Go 示例同样传入了二月二十九日,日期构造函数会将越界日期规范化,不能把它解释为有效的二月二十九日。
  • C++ 原例部分 $date 值为十位秒数,部分为十三位毫秒数,且 $match 下界为 1577836800、上界为 1609459200000。BSON 日期采用毫秒;若意图筛选 2020 年,下界应为 1577836800000,样本中其余秒数也应按对应时间改为毫秒。原例输出中 1970 年的日期和 TJ 的 1215 合计来自错误单位;正确的 2020 年筛选应排除值为 1024 的 2021 年订单,TJ 的两笔 2020 年订单合计为 191。
  • Python 模板把连接字符串的代码块标为 PHP,并链接 PHP 连接指南;Kotlin 连接字符串块标为 Java。这里仅修正代码块的显示语言,代码字符保持原样。Python 操作说明中的 Pipeline 名称与实际使用的 pipeline 列表也有差异,依照所示 Python API 构建阶段列表。
  • Java、Kotlin、PHP 与 Scala 样本的部分时间没有显式时区,Scala 的 SimpleDateFormat 默认时区也会影响解析。原文的不同语言输出存在时区或时间差异,不能把所有显示字符串直接视为相同的 UTC 时刻。
  • 模板不是可直接运行的最终程序。须补齐相关导入、集合引用、样本插入和阶段代码,再替换连接字符串;原有 API 签名与示例占位值均保留。Scala 原模板使用 Thread.sleep(1000) 配合异步订阅,固定等待不能保证操作完成,应在实际应用中明确协调完成顺序。

来源与许可

原文:Group and Total Data,作者与版权归属 MongoDB Documentation Team / MongoDB, Inc.。本文将原文叙事从英文译为中文,并按语言展开原页面的选项卡;源码、英文注释与示例输出保持原样,规范修订列于上文。示例输出来自原文,并非译者重新执行所得。

原文文档依据 Creative Commons Attribution-NonCommercial-ShareAlike 3.0 Unported 提供,许可关联见同版本官方 README。本译文采用相同许可。你可以在遵守许可的前提下复制、分享和改编:保留合理署名及原文与许可链接,明确注明翻译或其他修改,仅作非商业用途,并以相同或该许可证准许的兼容后续许可分享改编作品。不得施加限制他人行使许可权利的额外条款或有效技术限制。

许可不提供担保;其他权利或个别材料的限制仍可能适用。完整权利、义务、免责和终止规定以许可证法律文本为准。许可并不自动授权 MongoDB 商标或认可本译文。

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容