Skip to content

02. Why Machine Learning?

ML exists because some problems are too complex, too variable, or too data-rich for humans to write explicit rules for.

This page explains the specific problem categories where ML wins — and why the world’s most valuable companies depend on it.


Imagine trying to write rules for:

  • “Is this email spam?” — Spammers evolve daily. 10,000+ rule variations needed.
  • “Is this photo a cat?” — Impossible to describe every pixel pattern for every cat.
  • “What movie should I recommend?” — 300 million users × 10,000 movies = unique preferences.

Rules don’t scale. ML does.


mindmap
root((Why ML?))
Prediction
House prices
Stock trends
Disease risk
Classification
Spam or not
Fraud or not
Cat or dog
Recommendation
Netflix
Amazon
Spotify
Pattern Recognition
Face ID
Voice assistant
Medical imaging
Forecasting
Weather
Demand planning
Traffic
Decision Making
Loan approval
Ad targeting
Route optimization

Problem: Given past data, what will happen next?

Without ML: Someone manually studies trends and guesses.

With ML: Train on historical data → model predicts future values.

ExampleInput FeaturesPrediction
House pricesize, location, roomsPrice ($)
Disease riskage, vitals, historyProbability of condition
Sales forecastpast sales, season, eventsNext month revenue

Problem: Given an input, which category does it belong to?

Without ML: Write rules for every category boundary.

With ML: Learn boundaries from thousands of labeled examples.

# Spam classification - learned, not hand-coded
model.predict("You have won $1,000,000!") # → spam (98% confident)
model.predict("Call me when you're free") # → not_spam (99% confident)

Problem: Given a user and a catalog, what’s most relevant for this person right now?

Without ML: Manually create “genres” and recommend similar items.

With ML: Learn user preferences from implicit behavior (clicks, watches, purchases).

Netflix impact: Their recommendation system is estimated to save $1 billion/year by reducing churn from users not finding content.


Problem: Find meaningful signals in complex, high-dimensional data.

ApplicationDataPattern Found
Face recognitionPixel arrays”This is Person A”
Speech-to-textAudio waveforms”This is the word ‘hello‘“
Medical imagingX-ray pixels”This looks like a tumor”
Credit card fraudTransaction sequence”This pattern = fraud”

Human rules cannot describe what “a face” looks like in pixels. ML learns it from millions of examples.


Real-World Scale: Why Major Companies Depend on ML

Section titled “Real-World Scale: Why Major Companies Depend on ML”
flowchart LR
A[Your watch history] --> B[ML Model]
C[300M users' behavior] --> B
D[10k+ titles] --> B
B --> E[Personalized top 20]
E --> F[You keep watching]
F --> G[No churn = $1B saved/yr]
  • Product recommendation: 35% of revenue comes from ML-driven suggestions
  • Demand forecasting: warehouses stock the right items in the right locations
  • Fraud detection: real-time scoring of every transaction
  • ETA prediction: learned from millions of trips
  • Surge pricing: demand/supply forecasting in real time
  • Driver-rider matching: optimization at city scale
  • Discover Weekly: generated fresh every Monday for 600M users
  • Audio analysis: ML identifies tempo, key, energy, danceability of every song
  • Podcast recommendations: NLP on podcast transcripts
  • Search ranking: hundreds of ML models blend 200+ signals
  • Google Translate: neural machine translation for 100+ languages
  • Gmail spam: blocks 15 billion spam emails per day

ML works best when all three are true:

flowchart TD
A[✓ Lots of data] --> D[ML Works]
B[✓ Clear inputs/outputs] --> D
C[✓ Patterns exist in data] --> D

Without data → can’t learn. Without clear I/O → can’t train. Without patterns → nothing to learn.


SituationUse RulesUse ML
Calculate tax (fixed formula)✓
Detect spam✓
Validate an email address✓
Recommend movies✓
Business rule: discount > 20% needs approval✓
Predict customer churn✓
Sort a list✓
Recognize faces✓

# Demonstrating why ML beats rules for classification
# Rules approach - fragile
def spam_rules(email):
keywords = ["free money", "click here", "you won", "limited offer"]
return any(kw in email.lower() for kw in keywords)
# ML approach - learned from data
from sklearn.naive_bayes import MultinomialNB
from sklearn.feature_extraction.text import CountVectorizer
emails = [
("Get free money now", 1),
("Meeting notes attached", 0),
("You've been selected!", 1),
("Project deadline Friday", 0),
# ... thousands more in practice
]
texts, labels = zip(*emails)
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(texts)
model = MultinomialNB()
model.fit(X, labels)
# Handles novel phrasing rules can't catch
new_email = "Claim your exclusive reward today"
print(model.predict(vectorizer.transform([new_email]))) # → [1] spam

Q: Why do major tech companies invest so heavily in ML?

A: Because ML solves problems that don’t scale with human effort. Netflix can’t manually curate recommendations for 300M people. Amazon can’t write rules to predict demand for 350M products across thousands of locations. Google can’t hand-rank 8.5 billion daily searches. ML lets these companies personalize, predict, and optimize at a scale no team of humans could achieve with rules.


Q: What are the three conditions where ML outperforms rules?

A: (1) Large data exists — ML needs examples to learn from; more data = better patterns. (2) The problem has too many edge cases for rules — spam, fraud, and recommendations evolve constantly. (3) Input is unstructured — images, audio, and text can’t be described by simple rules; ML learns representations from raw data.


  • Using ML when a lookup table would do
  • Underestimating data requirements
  • Assuming ML “understands” — it finds correlations, not causes
  • Not measuring whether ML actually outperforms the baseline rules

ML exists to solve problems where:

  • Rules are too complex or constantly changing
  • Personalization at scale is required
  • Input is unstructured (images, audio, text)
  • Predictions improve with more data

Every major tech company uses ML at the core of its most valuable products.


← Previous: 01. What is ML? Next →: 03. ML Workflow