Writing Code With AI: A Realistic Workflow

The honest version of AI-assisted software development: test-driven prompts, small verifiable diffs, compiler guardrails, and how to stop fighting your tools.

July 9, 2026 | Noah Adeyemi Noah Adeyemi | 7 min read | 199 views
Writing Code With AI: A Realistic Workflow

Social media is filled with viral demos of developers typing three-word prompts and watching an artificial intelligence supposedly build a complete full-stack SaaS in twenty minutes. But inside engineering teams operating mission-critical software, those demos are recognized for what they are: party tricks. Generating code is trivial; maintaining, verifying, and refactoring code under real-world production traffic is where engineering actually happens. Here is the honest, battle-tested playbook for coding with AI without drowning in technical debt.

The Mental Model: The Supercharged Junior Engineer

The most dangerous mistake an engineer can make is treating a large language model like an all-knowing senior architect. A foundation model possesses photographic knowledge of syntax, libraries, and design patterns, but it has zero conceptual common sense and no situational memory of your company’s 2:00 AM outages.

Treat your AI coding assistant like an exceptionally eager, infinitely energetic junior engineer who:

  • Types at 800 words per minute without complaining.
  • Has read every Stack Overflow question and MDN doc ever published.
  • Will hallucinate a plausible-sounding function parameter the moment it gets confused.
  • Needs explicit constraints, test assertions, and strict peer review on every single pull request.

When you adopt this mental model, your role shifts from a typist to a technical director and code reviewer. The bottleneck of modern software engineering is no longer writing the lines of code—it is formulating unambiguous specifications.

Coding Modality Autonomy Level Feedback Loop Verification Burden Ideal Use Case
Ghost Autocomplete (Copilot/Supermaven) Micro (1-3 lines) Sub-second (<100ms) Immediate visual glance Repetitive boilerplate, standard loops, enum mapping
Editor In-Line Diff (Cursor / Windsurf) Function (20-60 lines) Fast (2-5 seconds) Git diff review per chunk Feature refactoring, API integration, regex parsing
Terminal Agent (Claude Code / Aider) Multi-file (100-500 lines) Moderate (15-60 seconds) Automated test suites & CI runs Cross-file migrations, lint fixes, test writing
Autonomous Cloud Agent (Devin-style) Repository-wide Slow (5-30 minutes) Comprehensive PR audit & security review Dependency upgrades, legacy tech debt remediation

The 4-Step Production Workflow

Over the past year of deploying AI tooling across dozens of production codebases, we refined a four-step discipline that maximizes velocity while keeping defect rates at zero:

Step 1: Test-Driven Generation (TDD as the Ultimate Prompt)

Writing prose prompts explaining how a complex function should behave is imprecise. English is full of ambiguity. Code is not.

The most effective prompt you can ever provide an AI is a failing test suite. Write your unit tests first—complete with edge cases, null checks, boundary numbers, and expected exception throws. Then prompt the model: "Write the implementation of this service class to make all tests in tests/Unit/PaymentTaxCalculatorTest.php pass cleanly without modifying the test file." The model has a mathematically verifiable target and will iterate until every assertion turns green.

Step 2: Enforce Small, Verifiable Diffs

Never ask an AI to "rewrite the user authentication service." An instruction that broad invites the model to rewrite architectural conventions, delete necessary error handling, and introduce subtle security flaws.

Restrict AI requests to changes small enough that you can hold the entire git diff in your head during review. If a diff exceeds 60 lines, break it into two sequential prompts. If you cannot review the code in one glance, you cannot guarantee its safety in production.

Step 3: Automated Guardrails Beat Human Willpower

You should never rely on your own eyes to catch whether an AI introduced a type mismatch or forgotten import. Build uncompromising guardrails into your local development loop:

  • Strict Static Analysis: Run PHPStan (level 8+), TypeScript in strict mode, or Pyright on every save. Static analysis catches hallucinated object methods before the code ever runs.
  • Pre-Commit AST Linters: Configure Husky or Git hooks to format code (Prettier, Pint, Ruff) and run type checks before allowing a commit.
  • Automated Mutation Testing: Use tools like Infection or Stryker to verify that your test suite genuinely asserts behavior rather than providing superficial line coverage.

Step 4: Architectural Isolation via Interfaces

Wrap AI-generated logic behind strict, typed interfaces. If you are having an AI generate a complex currency conversion algorithm or PDF generator, define the Interface contract yourself:

namespace App\Contracts;

interface CurrencyExchangeRateProvider
{
    /**
     * @param string $sourceCurrency ISO-4217 3-letter code
     * @param string $targetCurrency ISO-4217 3-letter code
     * @return float Exchange rate multiplier
     * @throws \App\Exceptions\ExchangeRateNotFoundException
     */
    public function getRate(string $sourceCurrency, string $targetCurrency): float;
}

By controlling the contract, you ensure that the AI's output is contained inside an isolated adapter. If the generated class ever develops bugs, your core domain entities remain untouched, and the class can be rewritten or replaced with zero collateral damage.

Practical Walkthrough: TDD as a Prompt Contract

Consider a real-world billing scenario: calculating prorated refunds for customer seat additions. If you prompt an LLM in English: "Write a function to prorate seat additions mid-cycle," it will almost certainly forget leap years, daylight saving time offsets, negative balance adjustments, or minimum billing thresholds.

Instead, you provide the model with a strict test specification:

it('calculates prorated seat additions accurately across month boundaries', function () {
    $calculator = new ProrationCalculator();
    
    // Cycle: 30 days total ($30/seat). Added on Day 20 (10 days remaining)
    $amount = $calculator->calculate(
        seatPrice: 30.00,
        cycleStartDate: Carbon::parse('2026-09-01'),
        cycleEndDate: Carbon::parse('2026-10-01'),
        changeDate: Carbon::parse('2026-09-21'),
        seatsAdded: 2
    );

    // Expect exactly 10 days * ($30/30 days) * 2 seats = $20.00
    expect($amount)->toEqual(20.00);
});

it('throws an exception if change date falls outside the active billing cycle', function () {
    $calculator = new ProrationCalculator();
    
    expect(fn () => $calculator->calculate(
        seatPrice: 30.00,
        cycleStartDate: Carbon::parse('2026-09-01'),
        cycleEndDate: Carbon::parse('2026-10-01'),
        changeDate: Carbon::parse('2026-10-15'),
        seatsAdded: 1
    ))->toThrow(InvalidBillingCycleException::class);
});

By establishing the mathematical expectations and exception classes upfront in executable code, the LLM has zero ambiguity. It generates the pure calculation class with exact boundary handling, and your test runner gives you instantaneous binary confirmation that the generated code is production-ready.

"AI will not replace software engineers. But software engineers who master test-driven verification and architectural guardrails will rapidly replace those who still type every boilerplate loop by hand."

— Noah Adeyemi, The Indox AI

Frequently Asked Questions

Key clarifications and practical answers addressed by The Indox editorial board.

Which IDE assistant is currently best for production development?

For interactive full-codebase editing, Cursor (and competitors like Windsurf) leads the pack because it indexes your entire AST and embeds symbol-aware context directly into its composer agent. For terminal-centric and command-line workflows, tools like Claude Code and Aider offer unparalleled git-integrated execution.

How do I prevent AI from introducing security vulnerabilities?

Never permit AI assistants to generate raw SQL strings or unescaped HTML interpolation. Always enforce parameterized prepared statements, strict CSRF validation, and run automated static application security testing (SAST) tools like Snyk or GitHub Dependabot in your CI pipeline.

Will AI code generators eliminate junior developer jobs?

The role of the junior developer is shifting from writing syntax to reading syntax. The junior engineers who flourish in the AI era are those who cultivate strong code reading skills, understand distributed systems principles, and learn how to debug systems when the automated tools fail.

The Pre-Commit AI Code Verification Checklist

Before you merge any pull request containing AI-assisted code into your main branch, run through this five-point sanity checklist:

  • Boundary Zero Check: Did you test with empty strings, negative integers, null values, and empty arrays? AI models routinely assume well-formed input.
  • Dependency Audit: Did the model import an existing library function, or did it hallucinate a non-existent helper function that passes locally only because of loose typing?
  • SQL & Injection Escaping: Are all database queries parameterized? Verify that string concatenation is never used to build queries.
  • Performance & Loop Nesting: Check that the model didn't introduce hidden O(N2) lookups (like querying a database or searching an array inside a nested foreach loop).
  • Documentation Freshness: Does the docblock comment accurately match the implementation? Models frequently write outdated docstrings when refactoring.

Final Takeaway

Coding with AI is neither a magic silver bullet nor a reckless gimmick. When deployed with disciplined test-driven prompts, rigorous QA test case verification, and uncompromising compiler assertions, it grants individual software engineers the velocity of a ten-person team without sacrificing software durability.

Master Architecture: Daily developer pair-programming workflows are cataloged in our comprehensive 2026 AI Software Engineering Playbook, analyzing developer latency, test-first scaffolding, and multi-agent coordination.

Tags: #Workflow #Coding #Developer Tools
Noah Adeyemi
Written By

Noah Adeyemi

Noah Adeyemi is a systems architect and quality engineering lead with over a decade of experience designing fault-tolerant distributed pipelines, CI/CD test automation harnesses, and high-concurrency microservices. Before joining The Indox AI as Lead QA Editor, Noah led test infrastructure teams across fintech and developer platform startups, where he spearheaded deterministic contract-testing frameworks and model-evaluation pipelines. At The Indox, Noah directs empirical benchmarking for AI code generation, agentic coding tools, and LLM test compilation, turning ambiguous agile requirements into rigorous, reproducible engineering assets.

Discussion (0)

No comments yet. Be the first to start the discussion!

Leave a Comment

Your email address will not be published. Required fields are marked *

The Indox AI Newsletter

Ideas That Help You Build Smarter with AI.

Calm, high-signal writing delivered to your inbox every week. Deep dives into LLM performance benchmarks, agent architectures, and hands-on engineering workflows.

Continue Reading

Related Articles