python - How to hash PySpark DataFrame to get a float returned?

Question

asked Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

Let's say I have spark dataframe

+--------+-----+
|  letter|count|
+--------+-----+
|       a|    2|
|       b|    2|
|       c|    1|
+--------+-----+

Then I wanted to find mean. So, I did

df = df.groupBy().mean('letter')

which give a dataframe

+------------------+
|       avg(letter)|
+------------------+
|1.6666666666666667|
+------------------+

how can I hash it to get only value 1.6666666666666667 like df["avg(letter)"][0] in Pandas dataframe? Or any workaround to get 1.6666666666666667

Note: I need a float returned. Not a list nor dataframe.

Thank you

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Answer

深蓝 · Answer 1 · 2021-10-23T18:42:30+0000

answered Oct 24, 2021 by 深蓝 (71.8m points)

Take first:

>>> df.groupBy().mean('letter').first()[0]

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…