如何使用 Python 识别序列中最常出现的项?
问题
您需要识别序列中最常出现的项。
解决方案
我们可以使用计数器来跟踪序列中的项。
什么是计数器?
"计数器"是一种映射,它保存每个键的整数计数。更新现有键会增加其计数。此对象用于计算可哈希对象的实例或作为多集。
"计数器"是您进行数据分析时最好的朋友之一。
这个对象在 Python 中已经存在了相当长一段时间,所以对于你们中的很多人来说,这将是一个快速回顾。我们将从从集合导入 Counter 开始。
从集合导入 Counter
如果传统字典缺少键,则会引发键错误。如果找不到键,Python 的字典将以键错误回答。
# 一个空字典
dict = {}
# 检查空字典中的键
dict['mystring']
# 错误消息
---------------------------------------------------------------------------
KeyError Traceback (most recent call last)
<ipython-input-12-1e03564507c6> in <module>
3
4 # 检查空字典中的键
----> 5 dict['mystring']
6
7 # 错误消息
KeyError: 'mystring'
在这种情况下,我们如何避免键错误异常?
Counter 是字典的一个子类,具有非常类似于字典的行为,但是,如果您查找缺少的键而不是引发键错误,它只会返回零。
# 定义计数器 c = Counter()
# 检查不可用的键
print(f"Output\n{c['mystring']}")
输出
0
c['mystring'] += 1
print(f"Output\n{c}")
输出
Counter({'mystring': 1})
示例
print(f"Output\n{type(c)}")
输出
<class 'collections.Counter'>
序列中最常出现的项
计数器的另一个优点是,您可以列出对象列表,它会为您计数。它使我们无需构建循环来构建计数器。
Counter
('Peas porridge hot peas porridge cold peas porridge in the pot nine days old'.split())
输出
Counter({'Peas': 1,
'porridge': 3,
'hot': 1,
'peas': 2,
'cold': 1,
'in': 1,
'the': 1,
'pot': 1,
'nine': 1,
'days': 1,
'old': 1})
拆分将采取的方法是获取字符串并将其拆分为单词列表。它在空格处拆分。
"计数器"将循环遍历该列表并计算所有单词,并在输出中显示计数。
还有更多,我还可以计算短语中最常见的单词。
most_common() 方法将为我们提供经常出现的项目。
count = Counter('Peas porridge hot peas porridge cold peas porridge in the pot nine days old'.split())
print(f"Output\n{count.most_common(1)}")
输出
[('porridge', 3)]
Example
print(f"Output\n{count.most_common(2)}")
输出
[('porridge', 3), ('peas', 2)]
示例
print(f"Output\n{count.most_common(3)}")
输出
[('porridge', 3), ('peas', 2), ('Peas', 1)]
请注意,它返回了一个元组列表。元组的第一部分是单词,第二部分是其计数。
Counter 实例的一个鲜为人知的功能是,它们可以使用各种数学运算轻松组合。
string = 'Peas porridge hot peas porridge cold peas porridge in the pot nine days old' another_string = 'Peas peas hot peas peas peas cold peas' a = Counter(string.split()) b = Counter(another_string.split())
# Add counts
add = a + b
print(f"Output\n{add}")
输出
Counter({'peas': 7, 'porridge': 3, 'Peas': 2, 'hot': 2, 'cold': 2, 'in': 1, 'the': 1, 'pot': 1, 'nine': 1, 'days': 1, 'old': 1})
# Subtract counts
sub = a - b
print(f"Output\n{sub}")
输出
Counter({'porridge': 3, 'in': 1, 'the': 1, 'pot': 1, 'nine': 1, 'days': 1, 'old': 1})
最后,Counter 在将数据存储在容器中方面非常聪明。
如上所示,它在存储时将单词分组在一起,允许我们将它们放在一起,这通常称为多集。
我们可以使用元素一次提取一个单词。它不记住顺序,而是将短语中的所有单词放在一起。
示例
print(f"Output\n{list(a.elements())}")
输出
['Peas', 'porridge', 'porridge', 'porridge', 'hot', 'peas', 'peas', 'cold', 'in', 'the', 'pot', 'nine', 'days', 'old']
示例
print(f"Output\n{list(a.values())}")
输出
[1, 3, 1, 2, 1, 1, 1, 1, 1, 1, 1]
示例
print(f"Output\n{list(a.items())}")
输出
[('Peas', 1), ('porridge', 3), ('hot', 1), ('peas', 2), ('cold', 1), ('in', 1), ('the', 1), ('pot', 1), ('nine', 1), ('days', 1), ('old', 1)]
相关文章
有用资源
python 参考教程 - 该教程包含有关 python 的更多信息:https://www.cainiaomax.com/python/

