Writing Data in Apache Spark: A Complete Guide to Saving DataFrames

Learn how Spark writes data to files, tables, and databases—and what actually happens when you call write. in this tutorial you will learn writing data in apache spark


Apache Spark Learning Path

APACHE SPARK LEARNING PATH

✓ 1. What is Apache Spark?
✓ 2. Why Spark is Faster than Hadoop
✓ 3. Spark Architecture
✓ 4. Driver vs Executor
✓ 5. Cluster Managers
→ 6. RDD vs DataFrame vs Dataset
→ 7. SparkSession
✓ 8. Reading Data
✓ 9. Lazy Evaluation
✓ 10. DAG
✓ 11. Stages and Tasks
✓ 12. Shuffle
✓ 13. Narrow vs Wide Transformations
→ 14. Writing Data
→ 15. Partitioning
→ 16. Data Skew


Where Are We Now?

In the previous tutorials, we learned how Spark:

  • creates a Spark application,
  • uses the Driver and Executors,
  • reads data,
  • builds a DAG,
  • divides work into stages and tasks,
  • and handles shuffle and partition boundaries.

We have now reached the other side of the data pipeline:

Reading data gets information into Spark.

Writing data gets processed information out of Spark.

And writing is more interesting than it initially appears.


What You’ll Learn

By the end of this tutorial, you’ll understand:

  • How Spark writes DataFrames
  • DataFrameWriter
  • Writing CSV, JSON, Parquet and ORC
  • Writing to tables
  • Append vs overwrite
  • Save modes
  • save() vs format-specific methods
  • Writing partitioned data
  • Why Spark creates multiple output files
  • Why Spark output is usually a directory
  • How coalesce() and repartition() affect output files
  • Common production mistakes
  • Interview questions around Spark writes

1. How Does Spark Write Data?

Suppose we have a DataFrame:

df.show()
+---+-------+------+
|id |name |salary|
+---+-------+------+
|1 |Alice |80000 |
|2 |Bob |90000 |
|3 |Charlie|75000 |
+---+-------+------+

We want to save it to storage.

In PySpark, the main interface is:

df.write

This returns a DataFrameWriter.

The general pattern is:

df.write \
.format("parquet") \
.mode("overwrite") \
.save("/data/employees")

Conceptually:

DataFrame
│
▼
DataFrameWriter
│
├── format
├── mode
├── options
└── save
│
▼
Storage System

2. DataFrameWriter

DataFrameWriter provides the API Spark uses to write structured data.

For example:

df.write.parquet("/data/employees")

or:

df.write \
.format("parquet") \
.save("/data/employees")

These represent the same general operation.

You can configure:

DataFrameWriter
│
├── Format
│ ├── Parquet
│ ├── JSON
│ ├── CSV
│ ├── ORC
│ └── JDBC
│
├── Save Mode
│ ├── append
│ ├── overwrite
│ ├── ignore
│ └── error / errorifexists
│
├── Options
│ ├── header
│ ├── compression
│ ├── delimiter
│ └── partitioning
│
└── Destination
├── Filesystem
├── Object Storage
└── Database

3. Writing Parquet

Parquet is one of the most common formats for Spark workloads.

df.write.parquet("/data/employees")

You can also use:

df.write \
.format("parquet") \
.save("/data/employees")

A typical output might look like:

/data/employees/
│
├── part-00000-....snappy.parquet
├── part-00001-....snappy.parquet
├── part-00002-....snappy.parquet
└── _SUCCESS

Notice something important:

Spark does not normally create one single file.

It creates multiple part files.

We’ll come back to why.


4. Writing CSV

You can write a DataFrame as CSV:

df.write \
.option("header", True) \
.csv("/data/employees_csv")

Or:

df.write \
.format("csv") \
.option("header", True) \
.save("/data/employees_csv")

Output:

/data/employees_csv/
│
├── part-00000-....csv
├── part-00001-....csv
├── part-00002-....csv
└── _SUCCESS

Important

CSV is convenient for interoperability, but for large Spark pipelines, columnar formats such as Parquet or ORC are often preferable.


5. Writing JSON

df.write.json("/data/employees_json")

With options:

df.write \
.format("json") \
.option("compression", "gzip") \
.save("/data/employees_json")

Again, Spark writes multiple part files based on the DataFrame’s partitioning.


6. Writing ORC

Spark also supports ORC:

df.write.orc("/data/employees_orc")

Or:

df.write \
.format("orc") \
.save("/data/employees_orc")

ORC is another columnar storage format commonly used in big-data ecosystems.


7. save() vs Format-Specific Methods

Spark provides both generic and format-specific APIs.

Format-specific

df.write.parquet("/data/output")
df.write.csv("/data/output")
df.write.json("/data/output")

Generic

df.write \
.format("parquet") \
.save("/data/output")

The generic API becomes especially useful when the format is configurable.

For example:

output_format = "parquet"
df.write \
.format(output_format) \
.save("/data/output")

8. Save Modes

One of the most important parts of writing data is the save mode.

Spark provides four common save modes:

ModeBehavior
errorIfExistsFail if destination already exists
appendAdd new data
overwriteReplace existing data
ignoreDo nothing if destination exists

8.1 Error If Exists

This is the default behavior for many file-based writes.

df.write \
.mode("errorIfExists") \
.parquet("/data/employees")

If the destination already exists, Spark throws an error.

You may also see:

.mode("error")

9. Append Mode

Append adds new data to an existing destination.

df.write \
.mode("append") \
.parquet("/data/employees")

Conceptually:

Existing Data
│
├── part-00000
├── part-00001
│
▼
Append New Data
│
├── part-00002
├── part-00003
└── ...

This is useful for incremental pipelines.

For example:

Day 1 → Write January 1 data
Day 2 → Append January 2 data
Day 3 → Append January 3 data

Important

Append does not mean Spark opens existing Parquet files and physically adds rows inside them.

New files are generally created in the destination.


10. Overwrite Mode

Overwrite replaces the existing output.

df.write \
.mode("overwrite") \
.parquet("/data/employees")

This is common for full-refresh pipelines.

For example:

Source
│
▼
Transform
│
▼
Complete Dataset
│
▼
Overwrite Target

A typical use case:

Every night:
Read source
↓
Recalculate dataset
↓
Overwrite target

11. Ignore Mode

ignore tells Spark not to write if the destination already exists.

df.write \
.mode("ignore") \
.parquet("/data/employees")

If the destination doesn’t exist:

Write data

If it already exists:

Do nothing

This can be useful for workflows where an existing output should remain untouched.


12. Why Does Spark Create Multiple Files?

This is one of the most important concepts to understand.

Suppose your DataFrame has four partitions:

DataFrame
│
├── Partition 0
├── Partition 1
├── Partition 2
└── Partition 3

When Spark executes the write:

Executor 1
│
├── Partition 0 → part-00000
└── Partition 1 → part-00001
Executor 2
│
├── Partition 2 → part-00002
└── Partition 3 → part-00003

So:

4 partitions
↓
4 tasks
↓
approximately 4 output part files

This is a consequence of Spark’s distributed execution model.


13. Partitions and Output Files

A useful mental model is:

DataFrame
│
▼
Partitions
│
▼
Tasks
│
▼
Output Files

For many file-based writes:

Each task writes the data from its partition into an output file.

Therefore, if you have too many partitions, you can end up with too many small files.

For example:

10,000 partitions
↓
10,000 tasks
↓
Potentially thousands of output files

This is commonly called the small-files problem.


14. Controlling the Number of Output Files

Suppose we have:

df.rdd.getNumPartitions()

and get:

100

We can reduce the number of partitions:

df2 = df.coalesce(10)

Then write:

df2.write.parquet("/data/output")

This can reduce the number of output files.


15. coalesce() vs repartition()

This distinction becomes extremely important in production Spark jobs.

coalesce()

df.coalesce(10)

Typically reduces the number of partitions without requiring a full shuffle.

repartition()

df.repartition(10)

Typically causes a shuffle to redistribute the data.

Conceptually:

coalesce()
100 partitions
│
▼
10 partitions
less movement

versus:

repartition()
100 partitions
│
▼
SHUFFLE
│
▼
10 balanced partitions

So if your only goal is to reduce partitions before writing, coalesce() may be appropriate.

If you need redistribution or a different partitioning strategy, repartition() may be necessary.


16. Partitioned Writes

Spark can physically partition output data by one or more columns.

Suppose we have:

id | name | department | salary
---|---------|------------|-------
1 | Alice | IT | 80000
2 | Bob | HR | 70000
3 | Charlie | IT | 90000
4 | David | Finance | 85000

We can write:

df.write \
.partitionBy("department") \
.parquet("/data/employees")

Spark may create:

/data/employees/
│
├── department=IT/
│ ├── part-00000.parquet
│ └── part-00001.parquet
│
├── department=HR/
│ └── part-00000.parquet
│
└── department=Finance/
└── part-00000.parquet

This is partitioned storage.


17. Why Partition Data?

Partitioning can make queries significantly more efficient.

Suppose we query:

SELECT *
FROM employees
WHERE department = 'IT';

If the data is physically partitioned by department, Spark can potentially avoid scanning unrelated partitions.

Conceptually:

Without partitioning
employees/
├── file1
├── file2
├── file3
├── file4
└── file5
Query: department = IT
↓
Potentially scan many files

With partitioning:

employees/
├── department=IT/
├── department=HR/
└── department=Finance/
Query: department = IT
↓
Read IT partition

This is known as partition pruning.


18. Don’t Partition by High-Cardinality Columns

A common mistake is:

df.write \
.partitionBy("user_id") \
.parquet("/data/output")

If you have millions of users, this can create an enormous number of directories and files.

For example:

user_id=1/
user_id=2/
user_id=3/
...
user_id=10,000,000/

That’s usually a terrible storage layout.

Good partition columns are generally:

  • frequently filtered,
  • relatively low/moderate cardinality,
  • useful for data organization.

Examples often include:

date
year
month
country
region
department

The right choice depends on the workload.


19. Writing with Options

Spark allows format-specific options.

For CSV:

df.write \
.format("csv") \
.option("header", True) \
.option("delimiter", ",") \
.mode("overwrite") \
.save("/data/output")

For compression:

df.write \
.format("parquet") \
.option("compression", "snappy") \
.save("/data/output")

The available options depend on the data source and Spark version.


20. Compression

Compression reduces storage size and can reduce the amount of data transferred across storage/network systems.

For example:

df.write \
.option("compression", "snappy") \
.parquet("/data/output")

For analytics workloads, columnar formats plus compression can provide an effective combination:

DataFrame
│
▼
Columnar Format
│
▼
Compression
│
▼
Smaller Storage
│
└──→ Less I/O

The best compression choice depends on workload, storage, CPU, and downstream compatibility.


21. Writing to a Database

Spark can also write DataFrames to relational databases using JDBC.

Example:

df.write \
.format("jdbc") \
.option("url", "jdbc:postgresql://host:5432/mydb") \
.option("dbtable", "employees") \
.option("user", "username") \
.option("password", "password") \
.mode("append") \
.save()

The architecture becomes:

Spark DataFrame
│
▼
Executors
│
▼
JDBC
│
▼
Database

However, writing to a database is different from writing to distributed object storage.

You need to consider:

  • database connection limits,
  • transaction behavior,
  • parallel writes,
  • batch sizes,
  • target-table indexes,
  • database capacity.

Simply increasing Spark parallelism can overwhelm a database.


22. Is write an Action?

Yes.

Consider:

df2 = df.filter(df.salary > 80000)

This transformation is lazy.

Nothing has necessarily been executed yet.

But:

df2.write.parquet("/data/output")

requires Spark to execute the computation and produce the output.

Conceptually:

filter()
│
│ lazy
▼
Logical/Execution Plan
│
▼
write()
│
▼
Job Execution
│
├── Stages
├── Tasks
└── Output Files

So writing data is one of the points where the previously defined transformations must actually be executed.


23. What Happens Internally During a Write?

Suppose we run:

df.filter("salary > 80000") \
.groupBy("department") \
.count() \
.write \
.parquet("/data/output")

A simplified execution flow is:

DataFrame Operations
│
▼
Query Plan
│
▼
Optimized Plan
│
▼
Physical Plan
│
▼
Spark Job
│
▼
Stages
│
▼
Tasks
│
▼
Output Files

The groupBy() may introduce a shuffle.

Therefore:

Filter
│
▼
GroupBy
│
├── Shuffle
│
▼
Aggregation
│
▼
Write

This connects directly to the concepts covered earlier in the series:

Transformations → DAG → Stages → Shuffle → Tasks → Write


24. Writing Data to Object Storage

In real-world data engineering, Spark commonly writes to distributed/object storage such as:

Amazon S3
Azure Data Lake Storage
Google Cloud Storage
HDFS

For example:

df.write.parquet("s3a://my-bucket/warehouse/employees/")

The important architectural idea is:

                 Spark Cluster
              ┌─────────────────┐
              │                 │
              │  Driver         │
              │                 │
              │  Executors      │
              │   │   │   │     │
              └───┼───┼───┼─────┘
                  │   │   │
                  ▼   ▼   ▼
              Object Storage
                  │
        ┌─────────┼─────────┐
        ▼         ▼         ▼
      file1     file2     file3

Executors perform the distributed work and write output data to the destination.


25. A Complete Example

Let’s build a small pipeline.

Step 1 — Create DataFrame

from pyspark.sql import SparkSession
spark = SparkSession.builder \
.appName("WritingDataExample") \
.getOrCreate()
data = [
(1, "Alice", "IT", 80000),
(2, "Bob", "HR", 70000),
(3, "Charlie", "IT", 90000),
(4, "David", "Finance", 85000)
]
columns = ["id", "name", "department", "salary"]
df = spark.createDataFrame(data, columns)

Step 2 — Transform

result = df.filter(df.salary > 75000)

Step 3 — Write

result.write \
.mode("overwrite") \
.parquet("/data/employees")

Step 4 — Partitioned Write

result.write \
.mode("overwrite") \
.partitionBy("department") \
.parquet("/data/employees_partitioned")

26. Common Mistake: Expecting One Output File

A beginner might write:

df.write.parquet("/data/output")

and expect:

output.parquet

Instead, Spark generally produces:

output/
├── part-00000-....parquet
├── part-00001-....parquet
├── part-00002-....parquet
└── _SUCCESS

This is normal.

Spark is a distributed processing engine.

It is designed to write data in parallel.


27. Common Mistake: Using coalesce(1) Everywhere

You may see:

df.coalesce(1) \
.write \
.mode("overwrite") \
.csv("/data/output")

This can produce a single output partition/file.

But it also creates a major bottleneck for large datasets.

Instead of:

100 partitions
│
├── Executor
├── Executor
├── Executor
└── Executor

you effectively force:

100 partitions
│
▼
1 partition
│
▼
1 task
│
▼
1 output file

For a large dataset, that can destroy parallelism.

Use it carefully.

A single output file may be reasonable for a small dataset or a specific downstream requirement, but it should not be the default strategy for large data.


28. Common Mistake: Too Many Small Files

Suppose:

1 TB dataset
+
10,000 partitions

You might end up with thousands of relatively small files.

This can cause:

  • metadata overhead,
  • slower file listing,
  • inefficient reads,
  • more scheduling overhead,
  • storage-system pressure.

A better approach is to think about:

Dataset Size
+
Partition Count
+
Partitioning Strategy
+
File Size

as one design problem.


29. Common Mistake: Partitioning by the Wrong Column

This:

.partitionBy("user_id")

can be problematic when user_id has extremely high cardinality.

Instead, a column such as:

.partitionBy("date")

may be more appropriate if most downstream queries filter by date.

The correct choice should come from query patterns, not simply from the columns available in the dataset.


30. Common Mistake: Blindly Using Overwrite

This:

.mode("overwrite")

is powerful—but dangerous.

If your destination contains important data, an incorrectly scoped overwrite can replace data you intended to preserve.

Before using overwrite in production, understand:

  • what path/table is being overwritten,
  • whether the write is partition-scoped,
  • how the underlying table format handles writes,
  • what happens on job failure,
  • whether atomicity or transactional guarantees are available.

31. Performance Tips

⚡ Performance Tip #1 — Choose the Right File Format

For analytical Spark workloads:

CSV
↓
Parquet / ORC

is often a useful direction.

Columnar formats can reduce the amount of data Spark needs to read for analytical queries.


⚡ Performance Tip #2 — Avoid Unnecessary Shuffles Before Writing

For example:

df.repartition(1)

forces redistribution of data.

Don’t do this simply because you want a small number of output files.

First understand your partitioning requirements.


⚡ Performance Tip #3 — Control Output Partitioning

If you have:

5000 partitions

for a relatively small dataset, blindly writing it may create thousands of small files.

Consider whether:

df.coalesce(...)

or another partitioning strategy makes sense before the write.


⚡ Performance Tip #4 — Partition Based on Query Patterns

If users frequently run:

WHERE date = '2026-09-09'

then organizing data around date partitions may be valuable.

Partitioning is not just a storage decision.

It is also a query-performance decision.


32. 💡 Going Deeper: File Count Is a Parallelism Decision

A useful way to think about Spark output is:

Partition Count
↓
Task Count
↓
Output Parallelism
↓
File Count

Therefore, the number of output files isn’t an arbitrary Spark behavior.

It is closely related to how the data is partitioned when the write executes.

This is why understanding partitions is essential before optimizing Spark writes.

And that leads naturally to the next tutorial:

Partitioning in Apache Spark.


33. Common Interview Questions

1. Why does Spark create multiple output files?

Because Spark processes data in partitions and executes tasks in parallel. File-based writes generally produce output files from those distributed tasks.


2. What is DataFrameWriter?

DataFrameWriter is the Spark API used to write DataFrames to supported data sources.

Example:

df.write.parquet("/data/output")

3. What are Spark save modes?

The common modes are:

append
overwrite
ignore
errorIfExists

4. What is the difference between append and overwrite?

Append adds new output to an existing destination.

Overwrite replaces the existing destination according to the semantics of the data source/table being written.


5. Why does coalesce(1) create one output file?

Because it reduces the DataFrame to one partition, meaning the write has one partition’s worth of output work.


6. Why shouldn’t we use coalesce(1) for large datasets?

It can eliminate parallelism and create a single-task bottleneck.


7. What is partitionBy()?

It organizes output data into directory structures based on one or more columns.

Example:

df.write \
.partitionBy("year", "month") \
.parquet("/data/output")

8. What is partition pruning?

Partition pruning allows Spark to avoid reading irrelevant physical partitions when query predicates match partition columns.


9. Is writing data an action in Spark?

Yes. A write requires Spark to execute the necessary computation to materialize the result.


10. What causes the small-files problem?

Common causes include:

  • excessive partition counts,
  • frequent incremental writes,
  • poor partitioning strategy,
  • very granular partition columns.

11. repartition() vs coalesce()?

repartition()
↓
redistributes data
↓
typically shuffle
coalesce()
↓
reduces partitions
↓
typically avoids a full shuffle

12. Why is Parquet commonly used with Spark?

Because it is a columnar format that supports efficient analytical reads, schema information, compression, and column-level access.


34. Final Mental Model

When you see:

df.write \
.mode("overwrite") \
.partitionBy("date") \
.parquet("/data/output")

break it down mentally:

df
│
├── DataFrame
│
▼
write
│
├── DataFrameWriter
│
├── mode = overwrite
│
├── partitionBy = date
│
├── format = parquet
│
└── destination = /data/output
│
▼
Spark Execution
│
▼
Partitions
│
▼
Tasks
│
▼
Output Files
│
▼
date=YYYY-MM-DD/

That’s the core of Spark writing.


Key Takeaways

  • df.write provides Spark’s DataFrame writing API.
  • Spark supports formats such as Parquet, ORC, CSV, and JSON, as well as databases through connectors such as JDBC.
  • Save modes include append, overwrite, ignore, and errorIfExists.
  • Spark generally writes multiple part files, not one file.
  • Output file count is strongly influenced by the number of partitions at write time.
  • coalesce() can reduce partitions without a full shuffle in common cases.
  • repartition() redistributes data and typically introduces a shuffle.
  • partitionBy() creates a partitioned directory layout.
  • Good partitioning can enable partition pruning.
  • High-cardinality partition columns can create a serious small-files problem.
  • A write triggers execution of the transformations required to produce the output.
  • Understanding Spark’s execution model is essential for designing efficient write pipelines.

Continue Learning Apache Spark

We’ve now covered both sides of a basic Spark data pipeline:

             Apache Spark

        ┌───────────────────┐
        │   Read Data       │ ✓
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │ Transform Data    │ ✓
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │ DAG / Stages /    │ ✓
        │ Tasks / Shuffle    │
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │   Write Data      │ ✓
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │   Storage         │
        └───────────────────┘

But there is one question we haven’t answered deeply enough:

How should Spark divide data into partitions in the first place?

That’s where Spark performance starts getting much more interesting.

Next Article

Partitioning in Apache Spark: How Spark Divides Data Across Executors

We’ll learn:

  • What a partition actually is
  • How Spark determines partition counts
  • repartition() vs coalesce()
  • Hash partitioning
  • Range partitioning
  • Partition sizing
  • Partition imbalance
  • How partitions affect tasks and executors
  • How poor partitioning leads to slow Spark jobs

Next → Partitioning in Apache Spark


Follow the Complete Apache Spark Series

Start from the beginning and build your understanding step by step:

Spark Fundamentals → Architecture → Data APIs → Execution → Shuffle → Writing → Partitioning → Skew → Join Optimization → Performance Tuning

The goal isn’t just to learn Spark syntax.

It’s to understand what Spark is actually doing when your code runs.

Leave a Reply

Discover more from Geeky Codes

Subscribe now to keep reading and get access to the full archive.

Continue reading