数据框:如何在Scala中分组/计数然后按计数排序

时间:2018-08-07 11:14:14

标签: scala apache-spark

我有一个包含数千行的数据框,我想要的是对分组并计数一列,然后按输出顺序排序:我所做的就是这样:

import org.apache.spark.sql.hive.HiveContext
import sqlContext.implicits._


val objHive = new HiveContext(sc)
val df = objHive.sql("select * from db.tb")
val df_count=df.groupBy("id").count().collect()
df_count.sort($"count".asc).show()

2 个答案:

答案 0 :(得分:7)

您可以按以下方式使用sortorderBy

val df_count = df.groupBy("id").count()

df_count.sort(desc("count")).show(false)

df_count.orderBy($"count".desc).show(false)

请勿使用collect(),因为它会将数据作为Array带到驱动程序。

希望这会有所帮助!

答案 1 :(得分:1)

//import the SparkSession which is the entry point for spark underlying API to access
 import org.apache.spark.sql.SparkSession
 import org.apache.spark.sql.functions._

 val pathOfFile="f:/alarms_files/"
//create session and hold it in spark variable
val spark=SparkSession.builder().appName("myApp").getOrCreate()
//read the file below API will return DataFrame of Row
var df=spark.read.format("csv").option("header","true").option("delimiter", "\t").load("file://"+pathOfFile+"db.tab")
//groupBY id column and take count of the column and order it by count of the column
    df=df.groupBy(df("id")).agg(count("*").as("columnCount")).orderBy("columnCount")
//for projecting the dataFrame it will show only top 20 records
    df.show
//for projecting more than 20 records  eg:
    df.show(50)