Learn how Spark writes data to files, tables, and databases—and what actually happens when you call write. in this tutorial you will learn writing data in apache spark
Apache Spark Learning Path
APACHE SPARK LEARNING PATH
✓ 1. What is Apache Spark?
✓ 2. Why Spark is Faster than Hadoop
✓ 3. Spark Architecture
✓ 4. Driver vs Executor
✓ 5. Cluster Managers
→ 6. RDD vs DataFrame vs Dataset
→ 7. SparkSession
✓ 8. Reading Data
✓ 9. Lazy Evaluation
✓ 10. DAG
✓ 11. Stages and Tasks
✓ 12. Shuffle
✓ 13. Narrow vs Wide Transformations
→ 14. Writing Data
→ 15. Partitioning
→ 16. Data Skew
Where Are We Now?
In the previous tutorials, we learned how Spark:
- creates a Spark application,
- uses the Driver and Executors,
- reads data,
- builds a DAG,
- divides work into stages and tasks,
- and handles shuffle and partition boundaries.
We have now reached the other side of the data pipeline:
Reading data gets information into Spark.
Writing data gets processed information out of Spark.
And writing is more interesting than it initially appears.
What You’ll Learn
By the end of this tutorial, you’ll understand:
- How Spark writes DataFrames
DataFrameWriter- Writing CSV, JSON, Parquet and ORC
- Writing to tables
- Append vs overwrite
- Save modes
save()vs format-specific methods- Writing partitioned data
- Why Spark creates multiple output files
- Why Spark output is usually a directory
- How
coalesce()andrepartition()affect output files - Common production mistakes
- Interview questions around Spark writes
1. How Does Spark Write Data?
Suppose we have a DataFrame:
df.show()
+---+-------+------+|id |name |salary|+---+-------+------+|1 |Alice |80000 ||2 |Bob |90000 ||3 |Charlie|75000 |+---+-------+------+
We want to save it to storage.
In PySpark, the main interface is:
df.write
This returns a DataFrameWriter.
The general pattern is:
df.write \ .format("parquet") \ .mode("overwrite") \ .save("/data/employees")
Conceptually:
DataFrame │ ▼DataFrameWriter │ ├── format ├── mode ├── options └── save │ ▼Storage System
2. DataFrameWriter
DataFrameWriter provides the API Spark uses to write structured data.
For example:
df.write.parquet("/data/employees")
or:
df.write \ .format("parquet") \ .save("/data/employees")
These represent the same general operation.
You can configure:
DataFrameWriter│├── Format│ ├── Parquet│ ├── JSON│ ├── CSV│ ├── ORC│ └── JDBC│├── Save Mode│ ├── append│ ├── overwrite│ ├── ignore│ └── error / errorifexists│├── Options│ ├── header│ ├── compression│ ├── delimiter│ └── partitioning│└── Destination ├── Filesystem ├── Object Storage └── Database
3. Writing Parquet
Parquet is one of the most common formats for Spark workloads.
df.write.parquet("/data/employees")
You can also use:
df.write \ .format("parquet") \ .save("/data/employees")
A typical output might look like:
/data/employees/│├── part-00000-....snappy.parquet├── part-00001-....snappy.parquet├── part-00002-....snappy.parquet└── _SUCCESS
Notice something important:
Spark does not normally create one single file.
It creates multiple part files.
We’ll come back to why.
4. Writing CSV
You can write a DataFrame as CSV:
df.write \ .option("header", True) \ .csv("/data/employees_csv")
Or:
df.write \ .format("csv") \ .option("header", True) \ .save("/data/employees_csv")
Output:
/data/employees_csv/│├── part-00000-....csv├── part-00001-....csv├── part-00002-....csv└── _SUCCESS
Important
CSV is convenient for interoperability, but for large Spark pipelines, columnar formats such as Parquet or ORC are often preferable.
5. Writing JSON
df.write.json("/data/employees_json")
With options:
df.write \ .format("json") \ .option("compression", "gzip") \ .save("/data/employees_json")
Again, Spark writes multiple part files based on the DataFrame’s partitioning.
6. Writing ORC
Spark also supports ORC:
df.write.orc("/data/employees_orc")
Or:
df.write \ .format("orc") \ .save("/data/employees_orc")
ORC is another columnar storage format commonly used in big-data ecosystems.
7. save() vs Format-Specific Methods
Spark provides both generic and format-specific APIs.
Format-specific
df.write.parquet("/data/output")
df.write.csv("/data/output")
df.write.json("/data/output")
Generic
df.write \ .format("parquet") \ .save("/data/output")
The generic API becomes especially useful when the format is configurable.
For example:
output_format = "parquet"df.write \ .format(output_format) \ .save("/data/output")
8. Save Modes
One of the most important parts of writing data is the save mode.
Spark provides four common save modes:
| Mode | Behavior |
|---|---|
errorIfExists | Fail if destination already exists |
append | Add new data |
overwrite | Replace existing data |
ignore | Do nothing if destination exists |
8.1 Error If Exists
This is the default behavior for many file-based writes.
df.write \ .mode("errorIfExists") \ .parquet("/data/employees")
If the destination already exists, Spark throws an error.
You may also see:
.mode("error")
9. Append Mode
Append adds new data to an existing destination.
df.write \ .mode("append") \ .parquet("/data/employees")
Conceptually:
Existing Data │ ├── part-00000 ├── part-00001 │ ▼Append New Data │ ├── part-00002 ├── part-00003 └── ...
This is useful for incremental pipelines.
For example:
Day 1 → Write January 1 dataDay 2 → Append January 2 dataDay 3 → Append January 3 data
Important
Append does not mean Spark opens existing Parquet files and physically adds rows inside them.
New files are generally created in the destination.
10. Overwrite Mode
Overwrite replaces the existing output.
df.write \ .mode("overwrite") \ .parquet("/data/employees")
This is common for full-refresh pipelines.
For example:
Source │ ▼Transform │ ▼Complete Dataset │ ▼Overwrite Target
A typical use case:
Every night: Read source ↓ Recalculate dataset ↓ Overwrite target
11. Ignore Mode
ignore tells Spark not to write if the destination already exists.
df.write \ .mode("ignore") \ .parquet("/data/employees")
If the destination doesn’t exist:
Write data
If it already exists:
Do nothing
This can be useful for workflows where an existing output should remain untouched.
12. Why Does Spark Create Multiple Files?
This is one of the most important concepts to understand.
Suppose your DataFrame has four partitions:
DataFrame│├── Partition 0├── Partition 1├── Partition 2└── Partition 3
When Spark executes the write:
Executor 1 │ ├── Partition 0 → part-00000 └── Partition 1 → part-00001Executor 2 │ ├── Partition 2 → part-00002 └── Partition 3 → part-00003
So:
4 partitions ↓4 tasks ↓approximately 4 output part files
This is a consequence of Spark’s distributed execution model.
13. Partitions and Output Files
A useful mental model is:
DataFrame │ ▼Partitions │ ▼Tasks │ ▼Output Files
For many file-based writes:
Each task writes the data from its partition into an output file.
Therefore, if you have too many partitions, you can end up with too many small files.
For example:
10,000 partitions ↓10,000 tasks ↓Potentially thousands of output files
This is commonly called the small-files problem.
14. Controlling the Number of Output Files
Suppose we have:
df.rdd.getNumPartitions()
and get:
100
We can reduce the number of partitions:
df2 = df.coalesce(10)
Then write:
df2.write.parquet("/data/output")
This can reduce the number of output files.
15. coalesce() vs repartition()
This distinction becomes extremely important in production Spark jobs.
coalesce()
df.coalesce(10)
Typically reduces the number of partitions without requiring a full shuffle.
repartition()
df.repartition(10)
Typically causes a shuffle to redistribute the data.
Conceptually:
coalesce()100 partitions │ ▼10 partitions less movement
versus:
repartition()100 partitions │ ▼ SHUFFLE │ ▼10 balanced partitions
So if your only goal is to reduce partitions before writing, coalesce() may be appropriate.
If you need redistribution or a different partitioning strategy, repartition() may be necessary.
16. Partitioned Writes
Spark can physically partition output data by one or more columns.
Suppose we have:
id | name | department | salary---|---------|------------|-------1 | Alice | IT | 800002 | Bob | HR | 700003 | Charlie | IT | 900004 | David | Finance | 85000
We can write:
df.write \ .partitionBy("department") \ .parquet("/data/employees")
Spark may create:
/data/employees/│├── department=IT/│ ├── part-00000.parquet│ └── part-00001.parquet│├── department=HR/│ └── part-00000.parquet│└── department=Finance/ └── part-00000.parquet
This is partitioned storage.
17. Why Partition Data?
Partitioning can make queries significantly more efficient.
Suppose we query:
SELECT *FROM employeesWHERE department = 'IT';
If the data is physically partitioned by department, Spark can potentially avoid scanning unrelated partitions.
Conceptually:
Without partitioningemployees/├── file1├── file2├── file3├── file4└── file5Query: department = IT ↓Potentially scan many files
With partitioning:
employees/├── department=IT/├── department=HR/└── department=Finance/Query: department = IT ↓Read IT partition
This is known as partition pruning.
18. Don’t Partition by High-Cardinality Columns
A common mistake is:
df.write \ .partitionBy("user_id") \ .parquet("/data/output")
If you have millions of users, this can create an enormous number of directories and files.
For example:
user_id=1/user_id=2/user_id=3/...user_id=10,000,000/
That’s usually a terrible storage layout.
Good partition columns are generally:
- frequently filtered,
- relatively low/moderate cardinality,
- useful for data organization.
Examples often include:
dateyearmonthcountryregiondepartment
The right choice depends on the workload.
19. Writing with Options
Spark allows format-specific options.
For CSV:
df.write \ .format("csv") \ .option("header", True) \ .option("delimiter", ",") \ .mode("overwrite") \ .save("/data/output")
For compression:
df.write \ .format("parquet") \ .option("compression", "snappy") \ .save("/data/output")
The available options depend on the data source and Spark version.
20. Compression
Compression reduces storage size and can reduce the amount of data transferred across storage/network systems.
For example:
df.write \ .option("compression", "snappy") \ .parquet("/data/output")
For analytics workloads, columnar formats plus compression can provide an effective combination:
DataFrame │ ▼Columnar Format │ ▼Compression │ ▼Smaller Storage │ └──→ Less I/O
The best compression choice depends on workload, storage, CPU, and downstream compatibility.
21. Writing to a Database
Spark can also write DataFrames to relational databases using JDBC.
Example:
df.write \ .format("jdbc") \ .option("url", "jdbc:postgresql://host:5432/mydb") \ .option("dbtable", "employees") \ .option("user", "username") \ .option("password", "password") \ .mode("append") \ .save()
The architecture becomes:
Spark DataFrame │ ▼Executors │ ▼JDBC │ ▼Database
However, writing to a database is different from writing to distributed object storage.
You need to consider:
- database connection limits,
- transaction behavior,
- parallel writes,
- batch sizes,
- target-table indexes,
- database capacity.
Simply increasing Spark parallelism can overwhelm a database.
22. Is write an Action?
Yes.
Consider:
df2 = df.filter(df.salary > 80000)
This transformation is lazy.
Nothing has necessarily been executed yet.
But:
df2.write.parquet("/data/output")
requires Spark to execute the computation and produce the output.
Conceptually:
filter() │ │ lazy ▼Logical/Execution Plan │ ▼write() │ ▼Job Execution │ ├── Stages ├── Tasks └── Output Files
So writing data is one of the points where the previously defined transformations must actually be executed.
23. What Happens Internally During a Write?
Suppose we run:
df.filter("salary > 80000") \ .groupBy("department") \ .count() \ .write \ .parquet("/data/output")
A simplified execution flow is:
DataFrame Operations │ ▼ Query Plan │ ▼ Optimized Plan │ ▼ Physical Plan │ ▼ Spark Job │ ▼ Stages │ ▼ Tasks │ ▼ Output Files
The groupBy() may introduce a shuffle.
Therefore:
Filter │ ▼GroupBy │ ├── Shuffle │ ▼Aggregation │ ▼Write
This connects directly to the concepts covered earlier in the series:
Transformations → DAG → Stages → Shuffle → Tasks → Write
24. Writing Data to Object Storage
In real-world data engineering, Spark commonly writes to distributed/object storage such as:
Amazon S3Azure Data Lake StorageGoogle Cloud StorageHDFS
For example:
df.write.parquet("s3a://my-bucket/warehouse/employees/")
The important architectural idea is:
Spark Cluster
┌─────────────────┐
│ │
│ Driver │
│ │
│ Executors │
│ │ │ │ │
└───┼───┼───┼─────┘
│ │ │
▼ ▼ ▼
Object Storage
│
┌─────────┼─────────┐
▼ ▼ ▼
file1 file2 file3
Executors perform the distributed work and write output data to the destination.
25. A Complete Example
Let’s build a small pipeline.
Step 1 — Create DataFrame
from pyspark.sql import SparkSessionspark = SparkSession.builder \ .appName("WritingDataExample") \ .getOrCreate()data = [ (1, "Alice", "IT", 80000), (2, "Bob", "HR", 70000), (3, "Charlie", "IT", 90000), (4, "David", "Finance", 85000)]columns = ["id", "name", "department", "salary"]df = spark.createDataFrame(data, columns)
Step 2 — Transform
result = df.filter(df.salary > 75000)
Step 3 — Write
result.write \ .mode("overwrite") \ .parquet("/data/employees")
Step 4 — Partitioned Write
result.write \ .mode("overwrite") \ .partitionBy("department") \ .parquet("/data/employees_partitioned")
26. Common Mistake: Expecting One Output File
A beginner might write:
df.write.parquet("/data/output")
and expect:
output.parquet
Instead, Spark generally produces:
output/├── part-00000-....parquet├── part-00001-....parquet├── part-00002-....parquet└── _SUCCESS
This is normal.
Spark is a distributed processing engine.
It is designed to write data in parallel.
27. Common Mistake: Using coalesce(1) Everywhere
You may see:
df.coalesce(1) \ .write \ .mode("overwrite") \ .csv("/data/output")
This can produce a single output partition/file.
But it also creates a major bottleneck for large datasets.
Instead of:
100 partitions │ ├── Executor ├── Executor ├── Executor └── Executor
you effectively force:
100 partitions │ ▼ 1 partition │ ▼ 1 task │ ▼ 1 output file
For a large dataset, that can destroy parallelism.
Use it carefully.
A single output file may be reasonable for a small dataset or a specific downstream requirement, but it should not be the default strategy for large data.
28. Common Mistake: Too Many Small Files
Suppose:
1 TB dataset+10,000 partitions
You might end up with thousands of relatively small files.
This can cause:
- metadata overhead,
- slower file listing,
- inefficient reads,
- more scheduling overhead,
- storage-system pressure.
A better approach is to think about:
Dataset Size +Partition Count +Partitioning Strategy +File Size
as one design problem.
29. Common Mistake: Partitioning by the Wrong Column
This:
.partitionBy("user_id")
can be problematic when user_id has extremely high cardinality.
Instead, a column such as:
.partitionBy("date")
may be more appropriate if most downstream queries filter by date.
The correct choice should come from query patterns, not simply from the columns available in the dataset.
30. Common Mistake: Blindly Using Overwrite
This:
.mode("overwrite")
is powerful—but dangerous.
If your destination contains important data, an incorrectly scoped overwrite can replace data you intended to preserve.
Before using overwrite in production, understand:
- what path/table is being overwritten,
- whether the write is partition-scoped,
- how the underlying table format handles writes,
- what happens on job failure,
- whether atomicity or transactional guarantees are available.
31. Performance Tips
⚡ Performance Tip #1 — Choose the Right File Format
For analytical Spark workloads:
CSV ↓Parquet / ORC
is often a useful direction.
Columnar formats can reduce the amount of data Spark needs to read for analytical queries.
⚡ Performance Tip #2 — Avoid Unnecessary Shuffles Before Writing
For example:
df.repartition(1)
forces redistribution of data.
Don’t do this simply because you want a small number of output files.
First understand your partitioning requirements.
⚡ Performance Tip #3 — Control Output Partitioning
If you have:
5000 partitions
for a relatively small dataset, blindly writing it may create thousands of small files.
Consider whether:
df.coalesce(...)
or another partitioning strategy makes sense before the write.
⚡ Performance Tip #4 — Partition Based on Query Patterns
If users frequently run:
WHERE date = '2026-09-09'
then organizing data around date partitions may be valuable.
Partitioning is not just a storage decision.
It is also a query-performance decision.
32. 💡 Going Deeper: File Count Is a Parallelism Decision
A useful way to think about Spark output is:
Partition Count ↓Task Count ↓Output Parallelism ↓File Count
Therefore, the number of output files isn’t an arbitrary Spark behavior.
It is closely related to how the data is partitioned when the write executes.
This is why understanding partitions is essential before optimizing Spark writes.
And that leads naturally to the next tutorial:
Partitioning in Apache Spark.
33. Common Interview Questions
1. Why does Spark create multiple output files?
Because Spark processes data in partitions and executes tasks in parallel. File-based writes generally produce output files from those distributed tasks.
2. What is DataFrameWriter?
DataFrameWriter is the Spark API used to write DataFrames to supported data sources.
Example:
df.write.parquet("/data/output")
3. What are Spark save modes?
The common modes are:
appendoverwriteignoreerrorIfExists
4. What is the difference between append and overwrite?
Append adds new output to an existing destination.
Overwrite replaces the existing destination according to the semantics of the data source/table being written.
5. Why does coalesce(1) create one output file?
Because it reduces the DataFrame to one partition, meaning the write has one partition’s worth of output work.
6. Why shouldn’t we use coalesce(1) for large datasets?
It can eliminate parallelism and create a single-task bottleneck.
7. What is partitionBy()?
It organizes output data into directory structures based on one or more columns.
Example:
df.write \ .partitionBy("year", "month") \ .parquet("/data/output")
8. What is partition pruning?
Partition pruning allows Spark to avoid reading irrelevant physical partitions when query predicates match partition columns.
9. Is writing data an action in Spark?
Yes. A write requires Spark to execute the necessary computation to materialize the result.
10. What causes the small-files problem?
Common causes include:
- excessive partition counts,
- frequent incremental writes,
- poor partitioning strategy,
- very granular partition columns.
11. repartition() vs coalesce()?
repartition() ↓redistributes data ↓typically shufflecoalesce() ↓reduces partitions ↓typically avoids a full shuffle
12. Why is Parquet commonly used with Spark?
Because it is a columnar format that supports efficient analytical reads, schema information, compression, and column-level access.
34. Final Mental Model
When you see:
df.write \ .mode("overwrite") \ .partitionBy("date") \ .parquet("/data/output")
break it down mentally:
df│├── DataFrame│▼write│├── DataFrameWriter│├── mode = overwrite│├── partitionBy = date│├── format = parquet│└── destination = /data/output │ ▼ Spark Execution │ ▼ Partitions │ ▼ Tasks │ ▼ Output Files │ ▼ date=YYYY-MM-DD/
That’s the core of Spark writing.
Key Takeaways
df.writeprovides Spark’s DataFrame writing API.- Spark supports formats such as Parquet, ORC, CSV, and JSON, as well as databases through connectors such as JDBC.
- Save modes include append, overwrite, ignore, and errorIfExists.
- Spark generally writes multiple part files, not one file.
- Output file count is strongly influenced by the number of partitions at write time.
coalesce()can reduce partitions without a full shuffle in common cases.repartition()redistributes data and typically introduces a shuffle.partitionBy()creates a partitioned directory layout.- Good partitioning can enable partition pruning.
- High-cardinality partition columns can create a serious small-files problem.
- A write triggers execution of the transformations required to produce the output.
- Understanding Spark’s execution model is essential for designing efficient write pipelines.
Continue Learning Apache Spark
We’ve now covered both sides of a basic Spark data pipeline:
Apache Spark
┌───────────────────┐
│ Read Data │ ✓
└─────────┬─────────┘
│
▼
┌───────────────────┐
│ Transform Data │ ✓
└─────────┬─────────┘
│
▼
┌───────────────────┐
│ DAG / Stages / │ ✓
│ Tasks / Shuffle │
└─────────┬─────────┘
│
▼
┌───────────────────┐
│ Write Data │ ✓
└─────────┬─────────┘
│
▼
┌───────────────────┐
│ Storage │
└───────────────────┘
But there is one question we haven’t answered deeply enough:
How should Spark divide data into partitions in the first place?
That’s where Spark performance starts getting much more interesting.
Next Article
Partitioning in Apache Spark: How Spark Divides Data Across Executors
We’ll learn:
- What a partition actually is
- How Spark determines partition counts
repartition()vscoalesce()- Hash partitioning
- Range partitioning
- Partition sizing
- Partition imbalance
- How partitions affect tasks and executors
- How poor partitioning leads to slow Spark jobs
Next → Partitioning in Apache Spark
Follow the Complete Apache Spark Series
Start from the beginning and build your understanding step by step:
Spark Fundamentals → Architecture → Data APIs → Execution → Shuffle → Writing → Partitioning → Skew → Join Optimization → Performance Tuning
The goal isn’t just to learn Spark syntax.
It’s to understand what Spark is actually doing when your code runs.