python--pandas切片

作者: FTDdata | 来源:发表于2021-04-25 09:32 被阅读0次

python--pandas切片
python--pandas删除
python--pandas样式
15.Go_Slice(切片)
python--pandas分组聚合
python--pandas读取excel
Python的高级特性
切片
你能一口说出go中字符串转字节切片的容量嘛？
【golang】slice底层函数传参原理易错点

pandas的切片操作是python中数据框的基本操作，用来选择数据框的子集。

环境

python3.9
win10 64bit
pandas==1.2.1

准备数据

import pandas as pd
player_list = [[1,'M.S.Dhoni', 36, 75, 5428000],
               [2,'A.B.D Villers', 38, 74, 3428000],
               [3,'V.Kholi', 31, 70, 8428000],
               [4,'S.Smith', 34, 80, 4428000],
               [5,'C.Gayle', 40, 100, 4528000],
               [6,'J.Root', 33, 72, 7028000],
               [7,'K.Peterson', 42, 85, 2528000]]
index=pd.date_range('2020',periods=7,freq='D').map(lambda x:x.date().__str__())
df = pd.DataFrame(player_list, columns=['Id','Name', 'Age', 'Weight', 'Salary'],index=index)
print(df)

            Id           Name  Age  Weight   Salary
2020-01-01   1      M.S.Dhoni   36      75  5428000
2020-01-02   2  A.B.D Villers   38      74  3428000
2020-01-03   3        V.Kholi   31      70  8428000
2020-01-04   4        S.Smith   34      80  4428000
2020-01-05   5        C.Gayle   40     100  4528000
2020-01-06   6         J.Root   33      72  7028000
2020-01-07   7     K.Peterson   42      85  2528000

[]方法

可以用中括号[]完成对数据框的切片。利用列名对列进行切片，利用列的布尔序列对行进行切片。

# 选取Name列
print(df['Name'])

2020-01-01        M.S.Dhoni
2020-01-02    A.B.D Villers
2020-01-03          V.Kholi
2020-01-04          S.Smith
2020-01-05          C.Gayle
2020-01-06           J.Root
2020-01-07       K.Peterson
Name: Name, dtype: object

# 选取Name和Age列
print(df[['Name','Age']])

                     Name  Age
2020-01-01      M.S.Dhoni   36
2020-01-02  A.B.D Villers   38
2020-01-03        V.Kholi   31
2020-01-04        S.Smith   34
2020-01-05        C.Gayle   40
2020-01-06         J.Root   33
2020-01-07     K.Peterson   42

# 选取Name字段中包含字符o的行
print(df[df['Name'].str.contains('o')])

            Id        Name  Age  Weight   Salary
2020-01-01   1   M.S.Dhoni   36      75  5428000
2020-01-03   3     V.Kholi   31      70  8428000
2020-01-06   6      J.Root   33      72  7028000
2020-01-07   7  K.Peterson   42      85  2528000

# 多条件行切片
print(df[df['Name'].str.contains('o') & df['Age'].map(lambda x: x > 33)])

            Id        Name  Age  Weight   Salary
2020-01-01   1   M.S.Dhoni   36      75  5428000
2020-01-07   7  K.Peterson   42      85  2528000

iloc方法

用iloc方法，使用行列的位置对数据框进行切片。支持布尔切片。

行切片

只传入一个参数时，表示对行进行切片。参数为整数返回序列，参数为列表返回数据框。正数表示正向切片，
负数表示反向切片。

# 选取第一行（序列）
print(df.iloc[0])

Id                1
Name      M.S.Dhoni
Age              36
Weight           75
Salary      5428000
Name: 2020-01-01, dtype: object

# 选取第一行（数据框）
print(df.iloc[[0]])

            Id       Name  Age  Weight   Salary
2020-01-01   1  M.S.Dhoni   36      75  5428000

# 选取前2行
print(df.iloc[:2])

            Id           Name  Age  Weight   Salary
2020-01-01   1      M.S.Dhoni   36      75  5428000
2020-01-02   2  A.B.D Villers   38      74  3428000

# 选取第三行到行末
print(df.iloc[2:])

            Id        Name  Age  Weight   Salary
2020-01-03   3     V.Kholi   31      70  8428000
2020-01-04   4     S.Smith   34      80  4428000
2020-01-05   5     C.Gayle   40     100  4528000
2020-01-06   6      J.Root   33      72  7028000
2020-01-07   7  K.Peterson   42      85  2528000

# 选取1,3,5行(设置起止位置和步长)
print(df.iloc[:6:2])

            Id       Name  Age  Weight   Salary
2020-01-01   1  M.S.Dhoni   36      75  5428000
2020-01-03   3    V.Kholi   31      70  8428000
2020-01-05   5    C.Gayle   40     100  4528000

# 选取倒数第2行到行末
print(df.iloc[-2:])

            Id        Name  Age  Weight   Salary
2020-01-06   6      J.Root   33      72  7028000
2020-01-07   7  K.Peterson   42      85  2528000

# 选取4,5,6行（布尔列表切片）
print(df.iloc[[False, False, False, True, True, True, False]])

            Id     Name  Age  Weight   Salary
2020-01-04   4  S.Smith   34      80  4428000
2020-01-05   5  C.Gayle   40     100  4528000
2020-01-06   6   J.Root   33      72  7028000

# 选取Name字段中包含o字符的行
print(df.iloc[df['Name'].str.contains('o').to_list()])

            Id        Name  Age  Weight   Salary
2020-01-01   1   M.S.Dhoni   36      75  5428000
2020-01-03   3     V.Kholi   31      70  8428000
2020-01-06   6      J.Root   33      72  7028000
2020-01-07   7  K.Peterson   42      85  2528000

列切片

使用iloc方法进行列切片时，需要行参数设置为:,表示选取所有的行。列切片方法与行切片相同。

# 选取第一列(序列)
print(df.iloc[:, 0])

2020-01-01    1
2020-01-02    2
2020-01-03    3
2020-01-04    4
2020-01-05    5
2020-01-06    6
2020-01-07    7
Name: Id, dtype: int64

# 选取第一列(数据框)
print(df.iloc[:, [0]])

            Id
2020-01-01   1
2020-01-02   2
2020-01-03   3
2020-01-04   4
2020-01-05   5
2020-01-06   6
2020-01-07   7

# 选取列名中包含a的列
print(df.iloc[:, df.columns.str.contains('a')])

                     Name   Salary
2020-01-01      M.S.Dhoni  5428000
2020-01-02  A.B.D Villers  3428000
2020-01-03        V.Kholi  8428000
2020-01-04        S.Smith  4428000
2020-01-05        C.Gayle  4528000
2020-01-06         J.Root  7028000
2020-01-07     K.Peterson  2528000

组合切片

同时设置行参数与列参数，使用iloc进行组合切片。

# 选取第一行，第一列的元素
print(df.iloc[0, 0])

# 选取第1,3行，2,4列
print(df.iloc[[0, 2], [1, 3]])

                 Name  Weight
2020-01-01  M.S.Dhoni      75
2020-01-03    V.Kholi      70

# 选取Name中包含o的行，且列名中包含a的列
print(df.iloc[df['Name'].str.contains('o').to_list(), df.columns.str.contains('a')])

                  Name   Salary
2020-01-01   M.S.Dhoni  5428000
2020-01-03     V.Kholi  8428000
2020-01-06      J.Root  7028000
2020-01-07  K.Peterson  2528000

loc方法

使用loc方法，用行列的名字对数据框进行切片，同时支持布尔索引。

行切片

传入一个参数时，只对行进行切片。

# 选取索引为2020-01-01的行
print(df.loc['2020-01-01'])

Id                1
Name      M.S.Dhoni
Age              36
Weight           75
Salary      5428000
Name: 2020-01-01, dtype: object

# 选取索引为2020-01-01和2020-01-03的行
print(df.loc[['2020-01-01', '2020-01-03']])

            Id       Name  Age  Weight   Salary
2020-01-01   1  M.S.Dhoni   36      75  5428000
2020-01-03   3    V.Kholi   31      70  8428000

# 选取Name字段中包含o的行
print(df.loc[df['Name'].str.contains('o')])

            Id        Name  Age  Weight   Salary
2020-01-01   1   M.S.Dhoni   36      75  5428000
2020-01-03   3     V.Kholi   31      70  8428000
2020-01-06   6      J.Root   33      72  7028000
2020-01-07   7  K.Peterson   42      85  2528000

# 选取索引中包含3或4的行
print(df.loc[df.index.str.contains('3|4')])

            Id     Name  Age  Weight   Salary
2020-01-03   3  V.Kholi   31      70  8428000
2020-01-04   4  S.Smith   34      80  4428000

列切片

使用loc方法进行列切片时，行参数需要设置为:,表示选取所有行。列切片方法与行切片相同。

# 选取Name和Age列
print(df.loc[:, ['Name', 'Age']])

                     Name  Age
2020-01-01      M.S.Dhoni   36
2020-01-02  A.B.D Villers   38
2020-01-03        V.Kholi   31
2020-01-04        S.Smith   34
2020-01-05        C.Gayle   40
2020-01-06         J.Root   33
2020-01-07     K.Peterson   42

# 选取Name列及后面所有的列
print(df.loc[:, 'Name':])

                     Name  Age  Weight   Salary
2020-01-01      M.S.Dhoni   36      75  5428000
2020-01-02  A.B.D Villers   38      74  3428000
2020-01-03        V.Kholi   31      70  8428000
2020-01-04        S.Smith   34      80  4428000
2020-01-05        C.Gayle   40     100  4528000
2020-01-06         J.Root   33      72  7028000
2020-01-07     K.Peterson   42      85  2528000

# 选取包含a字符的所有列
print(df.loc[:, df.columns.str.contains('a')])

                     Name   Salary
2020-01-01      M.S.Dhoni  5428000
2020-01-02  A.B.D Villers  3428000
2020-01-03        V.Kholi  8428000
2020-01-04        S.Smith  4428000
2020-01-05        C.Gayle  4528000
2020-01-06         J.Root  7028000
2020-01-07     K.Peterson  2528000

组合切片

同时设置行参数和列参数，使用loc方法进行组合切片。

# 选取索引为2020-01-01和2020-01-03的行，且列名为Id和Name的列
print(df.loc[['2020-01-01', '2020-01-03'], ['Id', 'Name']])

            Id       Name
2020-01-01   1  M.S.Dhoni
2020-01-03   3    V.Kholi

# 选取Name中包含o的行，且列名中包含a的列
print(df.loc[df['Name'].str.contains('o'), df.columns.str.contains('a')])

                  Name   Salary
2020-01-01   M.S.Dhoni  5428000
2020-01-03     V.Kholi  8428000
2020-01-06      J.Root  7028000
2020-01-07  K.Peterson  2528000

# 根据索引位置与列名切片
print(df.loc[df.index[3:], ['Name', 'Weight']])

                  Name  Weight
2020-01-04     S.Smith      80
2020-01-05     C.Gayle     100
2020-01-06      J.Root      72
2020-01-07  K.Peterson      85

# 根据索引名与列位置切片
print(df.loc[['2020-01-01', '2020-01-03'], df.columns[[2, 4]]])

            Age   Salary
2020-01-01   36  5428000
2020-01-03   31  8428000

filter方法

filter方法与loc方法类似，都是基于索引名和列名进行切片。

# 选取列(默认)
print(df.filter(items=['Name', 'Weight']))

                     Name  Weight
2020-01-01      M.S.Dhoni      75
2020-01-02  A.B.D Villers      74
2020-01-03        V.Kholi      70
2020-01-04        S.Smith      80
2020-01-05        C.Gayle     100
2020-01-06         J.Root      72
2020-01-07     K.Peterson      85

# 选取行
print(df.filter(items=['2020-01-01', '2020-01-03'], axis=0))

            Id       Name  Age  Weight   Salary
2020-01-01   1  M.S.Dhoni   36      75  5428000
2020-01-03   3    V.Kholi   31      70  8428000

# 选取列名包含字符e的列
print(df.filter(like='e'))

                     Name  Age  Weight
2020-01-01      M.S.Dhoni   36      75
2020-01-02  A.B.D Villers   38      74
2020-01-03        V.Kholi   31      70
2020-01-04        S.Smith   34      80
2020-01-05        C.Gayle   40     100
2020-01-06         J.Root   33      72
2020-01-07     K.Peterson   42      85

# 选取列名是字符e结尾的列
print(df.filter(regex='e$'))

                     Name  Age
2020-01-01      M.S.Dhoni   36
2020-01-02  A.B.D Villers   38
2020-01-03        V.Kholi   31
2020-01-04        S.Smith   34
2020-01-05        C.Gayle   40
2020-01-06         J.Root   33
2020-01-07     K.Peterson   42

python--pandas切片
pandas的切片操作是python中数据框的基本操作，用来选择数据框的子集。环境 python3.9 win1...
python--pandas删除
删除是数据清洗中的高频操作，本文基于pandas，介绍其dataFrame的一些删除操作，包括了删除行，删除列，删...
python--pandas样式
pandas的DataFrame类拥有style属性，style属性返回Styler类，Styler类的apply...
15.Go_Slice(切片)
Go 切片定义切片切片初始化 len()和cap()函数空(nil)切片切片拦截 append() 和co...
python--pandas分组聚合
分组聚合是数据处理中常见的场景，在pandas中用groupby方法实现分组操作，用agg方法实现聚合操作。环境...
python--pandas读取excel
对excel文件的读取是数据分析中常见的，在python中，pandas库的read_excel方法能够读取exc...
Python的高级特性
切片 list切片 tuple切片 str切片迭代在Python中迭代是通过for ... in ...来实现...
切片
切片前的准备 1切片前，备份图片；切片后，也要保存文件，为日后修改做准备。 2切片前设置切片文件夹（？不懂）切片...
你能一口说出go中字符串转字节切片的容量嘛？
神奇的现象切片，切片，又是切片! 前一篇文章讲的是切片，今天遇到的神奇问题还是和切片有关，具体怎么个神奇...
【golang】slice底层函数传参原理易错点
切片底层结构切片的底层结构主要包括引用数组的地址data，切片长度len与切片容量cap。以切片为参数调用函数...

python--pandas切片

环境

准备数据

[]方法

iloc方法

行切片

列切片

组合切片

loc方法

行切片

列切片

组合切片

filter方法

相关文章

python--pandas切片

python--pandas删除

python--pandas样式

15.Go_Slice(切片)

python--pandas分组聚合

python--pandas读取excel

Python的高级特性

切片

你能一口说出go中字符串转字节切片的容量嘛？

【golang】slice底层函数传参原理易错点

网友评论

延伸阅读

深度阅读

栏目导航

热点阅读

Python

数据分析

数据科学