如何操作文本数据
-
将姓名中的所有字符转换为小写。
In [4]: titanic["Name"].str.lower() Out[4]: 0 braund, mr. owen harris 1 cumings, mrs. john bradley (florence briggs th... 2 heikkinen, miss laina 3 futrelle, mrs. jacques heath (lily may peel) 4 allen, mr. william henry ... 886 montvila, rev. juozas 887 graham, miss margaret edith 888 johnston, miss catherine helen "carrie" 889 behr, mr. karl howell 890 dooley, mr. patrick Name: Name, Length: 891, dtype: str要把
Name列中的每个字符串变成小写,先选取该列(参见数据选择教程),再通过str访问器调用lower方法。这样会逐元素转换每个字符串。
时间序列教程中的日期时间对象具有 dt 访问器;与此类似,使用 str 访问器可以调用一系列专门的字符串方法。这些方法的名称通常与处理单个字符串的 Python 内置方法相同,但会逐元素应用到列中的每个值(还记得逐元素计算吗?)。
-
提取逗号前面的部分,创建包含乘客姓氏的新列
Surname。In [5]: titanic["Name"].str.split(",") Out[5]: 0 [Braund, Mr. Owen Harris] 1 [Cumings, Mrs. John Bradley (Florence Briggs ... 2 [Heikkinen, Miss Laina] 3 [Futrelle, Mrs. Jacques Heath (Lily May Peel)] 4 [Allen, Mr. William Henry] ... 886 [Montvila, Rev. Juozas] 887 [Graham, Miss Margaret Edith] 888 [Johnston, Miss Catherine Helen "Carrie"] 889 [Behr, Mr. Karl Howell] 890 [Dooley, Mr. Patrick] Name: Name, Length: 891, dtype: object使用
Series.str.split()后,每个值都会返回一个包含两个元素的列表:第一个元素是逗号之前的部分,第二个是逗号之后的部分。In [6]: titanic["Surname"] = titanic["Name"].str.split(",").str.get(0) In [7]: titanic["Surname"] Out[7]: 0 Braund 1 Cumings 2 Heikkinen 3 Futrelle 4 Allen ... 886 Montvila 887 Graham 888 Johnston 889 Behr 890 Dooley Name: Surname, Length: 891, dtype: object我们只关心表示姓氏的第一部分(元素 0),因此可以再次使用
str访问器,调用Series.str.get()提取它。字符串操作可以串联起来,在一次表达式中组合多个步骤。
更多提取字符串片段的信息,请参阅用户指南中“拆分与替换字符串”的部分。
-
提取泰坦尼克号上伯爵夫人的乘客数据。
In [8]: titanic["Name"].str.contains("Countess") Out[8]: 0 False 1 False 2 False 3 False 4 False ... 886 False 887 False 888 False 889 False 890 False Name: Name, Length: 891, dtype: boolIn [9]: titanic[titanic["Name"].str.contains("Countess")] Out[9]: PassengerId Survived Pclass ... Cabin Embarked Surname 759 760 1 1 ... B77 S Rothes [1 rows x 13 columns]对她的故事感兴趣?可以查看维基百科。
Series.str.contains()会检查Name列的每个字符串是否包含Countess,并分别返回True(包含)或False(不包含)。可以使用这个结果进行条件(布尔)索引,筛选数据;数据子集教程介绍过这种方法。泰坦尼克号上只有一位伯爵夫人,因此结果只有一行。
注意
Series.str.contains() 和 Series.str.extract() 还接受正则表达式,支持更强的文本提取能力,但这超出了本教程的范围。
更多提取字符串片段的信息,请参阅用户指南中“字符串匹配与提取”的部分。
-
哪位泰坦尼克号乘客的姓名最长?
In [10]: titanic["Name"].str.len() Out[10]: 0 23 1 51 2 21 3 44 4 24 .. 886 21 887 27 888 39 889 21 890 19 Name: Name, Length: 891, dtype: int64首先计算
Name列中每个姓名的长度。字符串方法Series.str.len()会分别应用到每个姓名,即逐元素计算。In [11]: titanic["Name"].str.len().idxmax() Out[11]: 307接下来需要找出表中姓名长度最大的对应位置,最好是索引标签。
idxmax()正是用于这个目的的方法。它操作的是整数而非字符串,因此无需使用str。In [12]: titanic.loc[titanic["Name"].str.len().idxmax(), "Name"] Out[12]: 'Penasco y Castellana, Mrs. Victor de Satode (Maria Josefa Perez de Soto y Vallejo)'根据行索引标签
307和列名Name,可以使用数据子集教程介绍的loc进行选择。
-
将
Sex列中的male替换为M,将female替换为F。In [13]: titanic["Sex_short"] = titanic["Sex"].replace({"male": "M", "female": "F"}) In [14]: titanic["Sex_short"] Out[14]: 0 M 1 F 2 F 3 F 4 M .. 886 M 887 F 888 F 889 M 890 M Name: Sex_short, Length: 891, dtype: strreplace()虽然不是字符串方法,却能方便地按照映射或词汇表转换特定值。它需要使用字典定义{from: to}映射。
警告
字符串访问器也提供 replace() 方法,用于替换特定字符序列。但如果要映射多个完整的值,就需要写成:
titanic["Sex_short"] = titanic["Sex"].str.replace("female", "F")
titanic["Sex_short"] = titanic["Sex_short"].str.replace("male", "M")
这种方式很繁琐,而且容易出错。想想看(也可以自行尝试):如果把这两条语句的执行顺序调换,会发生什么?
请记住
-
通过
str访问器调用字符串方法。 -
字符串方法逐元素执行,可以配合条件索引使用。
-
replace可以方便地按给定字典转换值。
完整介绍请参阅用户指南中的“处理文本数据”页面。











暂无评论内容