如何使用正则表达式和数据类型选择多个 DataFrame 列

pythonserver side programmingprogramming更新于 2026/2/16 23:08:17

DataFrame 可能与电子表格或数据库中的行和列数据集进行比较。DataFrame 是一个 2D 对象。

好吧,对 1D 和 2D 术语感到困惑?

1D(系列)和 2D(DataFrame)之间的主要区别在于您需要输入的信息点数量才能到达任何单个数据点。如果您以 Series 为例,并希望提取一个值,则只需要一个参考点,即行索引。

与表格 (DataFrame) 相比,一个参考点不足以获取数据点,您需要行值和列值的交集。

以下代码片段显示了如何从 csv 文件创建 Pandas DataFrame。

.read_csv() 方法默认创建一个 DataFrame。您可以通过搜索电影从 kaggle.com 下载电影数据集。

"""Script : Create a Pandas DataFrame from a csv file."""
import pandas as pd
movies_dataset pd read_csv "https://raw.githubusercontent.com/sasankac/TestDataSet/master/movies_data.csv")
# 1 print the type of object
type(movies_dataset)
# 2 print the top 5 records in a tabular format
movies_dataset head(5)



budget
id
original_language
original_title
popularity
release_date
revenue
runtime
status
title
vote_average
vote_count
0
237000000
19995
en
Avatar
150.437577
10/12/2009
2787965087
162.0
Released
Avatar
7.2
11800
1
300000000
285
en
Pirates of the Caribbean: At World's End
139.082615
19/05/2007
961000000
169.0
Released
Pirates of the Caribbean: At World's End
6.9
4500
2
245000000
206647
en
Spectre
107.376788
26/10/2015
880674609
148.0
Released
Spectre
6.3
4466
3
250000000
49026
en
The Dark Knight Rises
112.312950
16/07/2012
1084939099
165.0
Released
The Dark Knight Rises
7.6
9106
4
260000000
49529
en
John Carter
43.926995
7/03/2012
284139100
132.0
Released
John Carter
6.1
2124

2.选择单个 DataFrame 列。 将列名作为字符串或列表传递给索引运算符将返回列值作为 Series 或 DataFrame。

如果我们传入一个带有列名的字符串,您将得到一个 Series 作为输出,但是,传递只有一个列名的列表将返回 DataFrame。我们将通过示例看到这一点。

# select the data as series movies_dataset["title"]


0 Avatar
1 Pirates of the Caribbean: At World's End
2 Spectre
3 The Dark Knight Rises
4 John Carter
...
4798 El Mariachi
4799 Newlyweds
4800 Signed, Sealed, Delivered
4801 Shanghai Calling
4802 My Date with Drew
Name: title, Length: 4803, dtype: object


# select the data as DataFrame movies_dataset[["title"]]



title
0
Avatar
1
Pirates of the Caribbean: At World's End
2
Spectre
3
The Dark Knight Rises
4
John Carter
...
...
4798
El Mariachi
4799
Newlyweds
4800
Signed, Sealed, Delivered
4801
Shanghai Calling
4802
My Date with Drew

3.选择多个 DataFrame 列。

# Multiple DataFrame columns movies_dataset[["title""runtime","vote_average","vote_count"]]



title
runtime
vote_average
vote_count
0
Avatar
162.0
7.2
11800
1
Pirates of the Caribbean: At World's End
169.0
6.9
4500
2
Spectre
148.0
6.3
4466
3
The Dark Knight Rises
165.0
7.6
9106
4
John Carter
132.0
6.1
2124
...
...
...
...
...
4798
El Mariachi
81.0
6.6
238
4799
Newlyweds
85.0
5.9
5
4800
Signed, Sealed, Delivered
120.0
7.0
6
4801
Shanghai Calling
98.0
5.7
7
4802
My Date with Drew
90.0
6.3
16

为避免代码可读性问题,我始终建议定义一个变量来将列的名称保存为列表,并使用列名,而不是在代码中指定多个列名。

columns=["title","runtime","vote_average","vote_count"]movies_dataset[columns]



title
runtime
vote_average
vote_count
0
Avatar
162.0
7.2
11800
1
Pirates of the Caribbean: At World's End
169.0
6.9
4500
2
Spectre
148.0
6.3
4466
3
The Dark Knight Rises
165.0
7.6
9106
4
John Carter
132.0
6.1
2124
...
...
...
...
...
4798
El Mariachi
81.0
6.6
238
4799
Newlyweds
85.0
5.9
5
4800
Signed, Sealed, Delivered
120.0
7.0
6
4801
Shanghai Calling
98.0
5.7
7
4802
My Date with Drew
90.0
6.3
16

4.按列名选择 DataFrame 列。

.filter() 方法

此方法使用字符串搜索和选择列非常方便。其工作方式与 SQL 中的 like %% 参数基本相同。请记住,.filter() 方法仅通过检查列名而不是实际数据值来选择列。

.filter() 方法支持三个可用于选择操作的参数。

.like
.regex
.items

like parameter takes a string and attempts to find the column names that contain this string somewhere in the column name.


# Select the column that have a column name like "title" movies_dataset.filter(like="title").head(5)



original_title
title
0
Avatar
Avatar
1
Pirates of the Caribbean: At World's End
Pirates of the Caribbean: At World's End
2
Spectre
Spectre
3
The Dark Knight Rises
The Dark Knight Rises
4
John Carter
John Carter

.regex – More flexible way of selecting the columns names using regualr expressions

# Select the columns that end with "t"
movies_dataset.filter(regex=t).head()



budget
vote_count
0
237000000
11800
1
300000000
4500
2
245000000
4466
3
250000000
9106
4
260000000
2124

.items – 重复将列名作为字符串或列表传递给索引运算符,但不会引发 KeyError

按数据类型划分的 DataFrame 列。

如果您有兴趣过滤和仅使用某些数据类型,则 .select_dtypes 方法适用于列数据类型。

同样,.select_dtypes 方法在其包含或排除参数中接受多种数据类型(通过列表)或单一数据类型(作为字符串),并返回仅包含这些给定数据类型的列的 DataFrame。

.include 参数包含具有指定数据类型的列,而 .exclude 将忽略具有指定数据类型的列。

让我们首先看看数据类型和具有这些数据类型的列数

movies_dataset=pd.read_csv("https://raw.githubusercontent.com/sasankac/TestDataSet/master/movies_data.csv"
movies_dataset.dtypes.value_counts()


object 5
int64 4
float64 3
dtype: int64


a) 从 pandas DataFrames 中过滤整数数据类型。

movies_dataset select_dtypes(include="int")head(3)


2
1
0


b). 从 pandas DataFrames 中选择整数和浮点数据类型。

您可以指定多种数据类型,如下表所示。

movies_dataset select_dtypes(include=["int64","float"]).head(3)



budget
id
popularity
revenue
runtime
vote_average
vote_count
0
237000000
19995
150.437577
2787965087
162.0
7.2
11800
1
300000000
285
139.082615
961000000
169.0
6.9
4500
2
245000000
206647
107.376788
880674609
148.0
6.3
4466

c) 好吧,如果您只想要所有数字数据类型,只需指定数字

movies_dataset select_dtypes(include=["number"]).head(3)



budget
id
popularity
revenue
runtime
vote_average
vote_count
0
237000000
19995
150.437577
2787965087
162.0
7.2
11800
1
300000000
285
139.082615
961000000
169.0
6.9
4500
2
245000000
206647
107.376788
880674609
148.0
6.3
4466

d). 从 pandas DataFrames 中排除某些数据类型。


movies_dataset select_dtypes(exclude=["object"]).head(3)



budget
id
popularity
revenue
runtime
vote_average
vote_count
0
237000000
19995
150.437577
2787965087
162.0
7.2
11800
1
300000000
285
139.082615
961000000
169.0
6.9
4500
2
245000000
206647
107.376788
880674609
148.0
6.3
4466


注意:- 没有字符串数据类型需要处理,pandas 将它们转换为对象,因此如果您遇到异常"TypeError:无法理解数据类型"string"",请将字符串替换为对象。


相关文章


有用资源