Skip to main content

14 WAYS OF WRITING A WORD COUNT PROGRAM USING APACHE SPARK

Many of the IT professionals who are working on developing the Spark Applications or the other Engineers / Fresh Graduates who are willing to learn Spark, they will / have to start with the famous WordCount program.

In this Article, we'll see in how many ways a Spark WordCount program can be written using Spark's RDD, DataFrame & Dataset APIs and Scala as the programming language.

Before we get into the implementation of WordCount Program in Apache Spark, let's have a look at What is Bigdata?

What is Big Data

Big Data is one of the leading & trending technology in the IT industry, where many organisation are migrating / planning to migrate their projects into it, which deals with huge amounts of data like TeraBytes & PetaBytes at scale. For more info ClickHere

When we talk about Big Data, we always hear about two popular Open Source technologies:
  • Apache Hadoop
  • Apache Spark

Apache Hadoop:

The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures. For more info ClickHere

Apache Spark:

Apache Spark is a unified analytics engine for large-scale data processing.

Characteristics of Apache Spark:

Speed  - Run workloads 100x faster.

Ease of Use - Write applications quickly in Java, Scala, Python, R, and SQL.

Generality - Combine SQL, streaming, and complex analytics.

Runs Everywhere - Spark runs on Hadoop, Apache Mesos, Kubernetes, standalone, or in the cloud. It can access diverse data sources.

For more info ClickHere

Prerequisites / Environment Setup:

  • Java (JDK) 1.8
  • IntelliJ
  • SBT 1.3.1
  • Scala 2.12.11
  • Spark Core & Spark SQL dependencies
build.sbt:


Creating new Object file:


Constants Class:


Initializing SparkSession Object:


Preparing Input Data:


Okiee, now we have the setup ready and we can start writing our Word Count Program.

Writing WordCount Program:

After a good amount of research and putting all my Spark experience together for writing this WordCount program  took around 3 days and I ended up with 14 different variations. Let's see each one of those one by one:

1. Using: RDD - FlatMap, Map & ReduceByKey


2. Using: RDD - FlatMap & CountByValue


3. Using: RDD - Using FlatMap & CountByKey


4. Using: RDD - FlatMap, GroupBy & Map


5. Using: RDD - FlatMap, Map & AggregateByKey


6. Using: RDD - FlatMap, Map & CombineByKey


7. Using: RDD - FlatMap, Map & FoldByKey


8. Using : RDD - MapPartitions, Map & ReduceByKey


9. Using: RDD - MapPartitions, Map & ReduceByKey


10. Using: DataFrame - Select, Explode, GroupBy & Count


11. Using: DataFrame - Transform, Select, Explode, GroupBy & Count


12 Using : DataFrame - TempTable & SQL Query


13. Using : Dataset - FlatMap, GroupBy & Count


14. Using : Dataset - Transform, FlatMap, GroupBy & Count


Main Method:


That's it. Most of them look like similar to each other but there is a lot of difference in terms of the implementation and also in performance. 

Full code is available at my GitHub Repo

Hope this article is helpful for all the professionals who wants to learn / improve their programming skills on Apache Spark.

If you have any questions / queries on the above implementations of WordCount program, please do comment and also feel free to add more implementations if you find anything in the comments section below.

Happy Coding...!!!

Thanks You.

Comments