Skip to content

14. Future of AI

AI has progressed faster in the last 5 years than in the previous 50. The 2017 transformer paper, followed by GPT-3 (2020) and ChatGPT (2022), triggered a step-change in capability and adoption.

Understanding the near-term trajectory helps you make better decisions about what to learn and build.


Models trained to “think step by step” before answering. OpenAI o1/o3, Google Gemini 2.0 Flash Thinking, Claude’s extended thinking mode.

Impact: Dramatically better performance on math, coding, logic — domains that need multi-step reasoning.


Models that process text + images + audio + video together.

  • GPT-4o, Gemini 1.5 Pro, Claude 3.5 already multimodal
  • Next: real-time voice interaction, video understanding, integrated document/code/image reasoning

LLMs that don’t just respond — they take actions:

  • Browse the web, run code, call APIs
  • Plan multi-step tasks autonomously
  • Use tools (Model Context Protocol — MCP)

This is the current frontier. The shift from “chatbot” to “agent that gets things done” is underway.


Smaller, efficient models running locally:

  • Apple Intelligence, Gemini Nano on Android
  • Privacy-preserving — data stays on device
  • No latency, no API costs

  • Every cloud provider building specialized AI hardware (TPUs, Trainium, H100 clusters)
  • Vector databases, AI observability, LLM evaluation tools becoming standard
  • “AI stack” emerging as a new layer in every software architecture

Definition: AI that can match human performance across any cognitive domain.

Current state: No consensus on definition, timeline, or whether it’s achievable. Leading researchers disagree on whether current approaches scale to AGI.

What’s clear: If something close to AGI is achieved, the economic and social disruption would be enormous. Most serious AI safety research is motivated by this scenario.


  • AI already generating drug candidates, protein structures
  • Future: AI that generates and tests hypotheses autonomously
  • Could compress decades of research into years

  • EU AI Act (2024) — risk-based regulation framework
  • US executive orders on AI safety
  • Global frameworks for high-risk AI systems (weapons, critical infrastructure, healthcare)

Engineers in the near future will need to understand regulatory requirements the same way they understand GDPR today.


RiskConcern
MisinformationLLMs make synthetic media and text indistinguishable from real
Economic disruptionJob displacement faster than retraining pipelines
Concentration of powerAI capabilities concentrated in a few companies
BioweaponsLLMs lowering barrier for biological weapon design
Autonomous weaponsAI in military decisions without human oversight
MisalignmentPowerful AI systems optimizing for the wrong objective

NowNear Future
Learn to use LLM APIsUnderstand and build AI agents
Use AI tools for productivityArchitect AI-native applications
RAG basicsProduction AI pipelines at scale
Understand promptingUnderstand fine-tuning and evaluation

The engineers who will be most valuable are those who understand the full stack: from how transformers work to how to build reliable production AI systems.


Q: What is AGI and how far away are we?

A: AGI (Artificial General Intelligence) refers to AI that matches or exceeds human cognitive performance across all domains — not just one specific task. Current AI is narrow: GPT-4 writes well but can’t learn to drive. AGI would seamlessly transfer knowledge across domains. Timelines are genuinely unknown — serious estimates range from 5 to 50+ years, and some researchers believe current approaches cannot reach AGI without fundamental new ideas.


Q: What are AI agents and why are they significant?

A: AI agents are LLM-powered systems that can take actions in the world — not just generate text. They can call APIs, run code, browse the web, manage files, and chain multi-step reasoning to complete complex tasks. They’re significant because they shift AI from passive information retrieval to active task execution. Protocols like MCP (Model Context Protocol) are standardizing how agents interact with external tools and services.