Stop Paying for Idle Servers: How I Built a Flask ML App That Costs Almost Nothing on AWS

Introduction

Imagine you’ve built an amazing house price prediction website(flask ML App) using Flask and a machine learning model.

The application works perfectly. Users enter details like location, area, number of bedrooms, and the model predicts the property’s price within milliseconds.

There’s just one problem.

Most of the day, nobody is using it.

Yet your AWS EC2 instance keeps running 24 hours a day, quietly burning money.

This is one of the most common mistakes developers make when deploying personal projects and machine learning applications.

In this article, we’ll explore:

  • Why EC2 isn’t always the right choice
  • How to automatically stop idle servers
  • Why waking an EC2 instance on incoming traffic doesn’t work well
  • Better AWS architectures that cost almost nothing when idle
  • Which hosting option to choose for your Flask application

The Traditional Architecture

Many developers deploy Flask applications like this:

Simple,Reliable,But expensive.

Even if only five users visit your website in an entire day, your EC2 instance continues running.

You’re paying for compute that isn’t doing any useful work.

Can We Automatically Stop an Idle EC2 Instance?

Yes.

AWS provides an elegant solution using CloudWatch and Lambda.

Architecture

CloudWatch continuously monitors metrics such as:

  • CPU Utilization
  • Network Traffic
  • Disk I/O

If the instance remains idle for a predefined period, a Lambda function simply calls:

StopInstances()

Your EC2 instance shuts down automatically.

No manual intervention required.

But Can EC2 Start Automatically When Someone Visits the Website?

This seems like the perfect solution.

Unfortunately… Not really. Let’s see what happens.

The entire process may take anywhere from 30 seconds to several minutes.

During that time:

  • Users experience long delays
  • Requests may fail
  • Search engines may mark the website as unavailable

This creates a poor user experience.

Why Traffic Cannot “Wake Up” a Stopped EC2 Instance

An EC2 instance that is stopped has:

  • No running web server
  • No Flask application
  • No Nginx
  • No Gunicorn
  • No process listening on port 80 or 443

Which means there is nothing capable of receiving the very first HTTP request.

It’s like calling a phone that has been switched off.

No one answers.

Better Solution #1 — AWS Lambda

If your machine learning model is reasonably small, Lambda is an excellent option.

Advantages

  • Pay only when requests arrive
  • Zero server management
  • Automatic scaling
  • Extremely low cost

This works beautifully for:

  • Scikit-learn
  • XGBoost
  • LightGBM
  • Small TensorFlow models

Better Solution #2 — AWS App Runner

If you already have a Flask application packaged inside Docker:

Benefits include:

  • HTTPS by default
  • Automatic deployments
  • Auto scaling
  • No EC2 management
  • Simple CI/CD pipeline

For many production Flask applications, this is the sweet spot.

Better Solution #3 — ECS Fargate

If your application is containerized:

Advantages:

  • No virtual machine management
  • Container-based deployment
  • Excellent scalability

Better Solution #4 — SageMaker Serverless Inference

If your primary goal is serving machine learning predictions:

Perfect for:

  • Deep learning models
  • Larger artifacts
  • Production inference workloads

Which Option Should You Choose?

Cost Comparison

My Recommendation

If I were deploying a house price prediction application today, my architecture would look like this:

If the application later grows to handle thousands of daily users, migrating to App Runner, ECS, or SageMaker becomes straightforward.

Key Takeaways

  • Don’t keep an EC2 instance running 24×7 for a website that receives very little traffic.
  • Automatically stopping idle EC2 instances is easy with CloudWatch and Lambda.
  • Automatically waking a stopped EC2 instance on the first web request creates unacceptable latency.
  • Serverless services like AWS Lambda and SageMaker Serverless are designed specifically for sporadic workloads.
  • For most Flask machine learning projects, AWS App Runner offers an excellent balance between simplicity, scalability, and cost.

Final Thoughts

Cloud computing isn’t just about deploying applications — it’s about deploying them efficiently.

As developers, we often focus on building models, APIs, and user interfaces while overlooking infrastructure costs. Choosing the right hosting strategy can reduce your cloud bill dramatically without sacrificing performance.

The next time you deploy a Flask machine learning application, ask yourself one question:

“Am I paying for compute, or am I paying for value?”

That simple question can save both money and operational complexity.

Leave a Reply

Discover more from Geeky Codes

Subscribe now to keep reading and get access to the full archive.

Continue reading