Just launched: 360° security audit to protect your legacy code from AI exploits.

Discover
Blog›Gen AI in Modernization

Legacy Code Modernization: A Practical Guide to Moving Code and Data Together

In This Article

TL;DR

  • Classify each application by friction and data coupling. Missing tests, missing documentation, and logic held in the database each add effort before transformation starts.
  • Decide where database-resident logic lands before transformation starts. Each stored procedure, trigger, and view stays in the database, moves into service code, or gets rewritten for the target database.
  • Sequence data changes with the code. Put data access behind one layer, change schemas in expand-and-contract steps, and move code in slices that each have a working data path.
  • Prove behavior across five layers. Compare API responses, database state, batch outputs, generated files, and downstream reports against a legacy baseline before cutover.
  • Estimate from drivers. Database logic, starting test coverage, dependency depth, integration count, and target-state clarity set the timeline.

What Legacy Code Modernization Involves

Legacy code modernization moves an application’s code to a current language, framework, or architecture, with the business running on it throughout. The code depends on a database that holds part of its logic, so the data layer moves in the same program. The work is complete when the modernized system produces the same results, data, and outputs as the legacy system.

Code migration moves code from one language, framework, or platform to another and keeps its behavior the same, as in Struts to Spring Boot or .NET Framework to .NET 8. A modernization program contains one or more code migrations, plus the architecture, data, and validation work around them.

Legacy is measured by friction. An application is legacy when the team can’t change it safely, because tests are missing, documentation is out of date, or the people who understood it have left.

Technical debt drives that friction. Forrester predicted that 75% of technology decision-makers would see their technical debt rise to a moderate or high level of severity by 2026 [1].

This guide covers distributed enterprise stacks, with examples drawn from five common legacy estates.

  • .NET Framework and ASP.NET WebForms to modern .NET.
  • Java EE and Struts to Spring Boot.
  • AngularJS to Angular or React.
  • VB6 and Delphi to .NET.
  • PL/SQL- and T-SQL-heavy applications, where the database holds much of the logic.

How to Choose a Code Modernization Path for Each Application

Code modernization runs one decision per application. Business value, rate of change, and data coupling decide the path, and the data-layer column below shows how much database work comes with each choice.

PathWhat ChangesData-Layer ImpactFits When
RetainNothing for nowNoneLow change rate on a supported runtime
RetireThe application is removedData is archived or merged elsewhere, and consumers are repointedFunctionality is duplicated or unused
RehostInfrastructureSame database, new network path and connection settingsA data center exit sets the deadline
ReplatformRuntime or managed services, with minimal code changeDriver and connection changes, or a new database engineA supported upgrade path exists
RefactorInternal structure, same behaviorData access layer cleaned up, schema largely unchangedThe domain model holds and tests can be built
RearchitectBoundaries, such as monolith to servicesSchema split by service, database logic redistributedScale or team ownership needs new boundaries
RebuildCode rewritten on a new stackData model redesigned, full data migrationThe stack has no upgrade path and requirements are recoverable
ReplaceA commercial or SaaS productData moved into the vendor’s schemaThe capability is a commodity

Retire and replace remove code from the estate, so a portfolio review that starts with those two paths shrinks the scope before any transformation work is priced.

Four criteria decide between refactoring and rebuilding an application.

CriterionPoints to RefactorPoints to Rebuild
Domain modelThe model matches how the business works todayThe business has outgrown the model
Upgrade pathThe framework has a supported path forwardThe stack has no upgrade path, as with VB6
TestabilityCharacterization tests can be built around the codeThe code is untestable in place, so tests wrap the legacy system’s inputs and outputs
Database couplingRules sit in application codeRules sit in stored procedures and need redesign in the new data model

The choice between refactoring and replatforming follows the same per-application logic.

Database Logic, Data Access, and Schema Dependencies in Legacy Code Modernization

Legacy applications keep part of their logic outside the application code. Three areas decide how much of a legacy code modernization program is data work.

  • Business logic in the database. Stored procedures, triggers, and views that hold business rules.
  • The data access layer. ORM mappings, embedded SQL, and transaction handling in the code.
  • Schema dependencies. Every service, report, and integration that reads a table.
Where business logic lives in a legacy application

Business Logic in Stored Procedures, Triggers, and Views

Three kinds of database object carry business rules in older applications.

  • Stored procedures. PL/SQL packages and T-SQL procedures hold pricing rules, eligibility checks, status transitions, and period-close calculations.
  • Triggers. They enforce rules on every write, including writes from batch jobs and integrations the application team doesn’t own.
  • Views. They encode the joins and filters that reports depend on.

Moving the application without deciding where that logic lands breaks behavior without a visible error. New code calls the same procedure with different parameters, or a trigger fires a second time once the service layer adds the same rule.

Each database object gets one of three decisions before transformation starts.

  • Keep in the database. Set-based logic that runs close to the data, such as bulk updates and period-close calculations. It stays on the same engine and gets characterization tests of its own.
  • Move into service code. Business rules that change often or need unit tests, such as pricing and eligibility. The procedure is retired once the service owns the rule.
  • Rewrite for the target database. Logic that stays in the database through an engine change, such as PL/SQL packages moving to PL/pgSQL.

Engine changes add their own conversion work, covered for an Oracle to PostgreSQL migration, a SQL Server to PostgreSQL migration, and a Db2 to PostgreSQL migration.

Data Access Layer Changes in a Code Migration

Every code migration touches the layer between the code and the database. ORM mappings, raw SQL strings, connection and transaction handling, and vendor-specific error codes all change when the framework or the database changes.

.NET Framework to modern .NET shows the scale. Microsoft describes EF Core as a total rewrite of Entity Framework with no direct upgrade path [2]. EF6 Code First migrations history doesn’t carry over, so teams start migrations fresh. EDMX models need conversion first, since EF Core doesn’t support the format.

Applications built on raw ADO.NET add an ORM decision, covered in an ADO.NET to EF Core migration.

Java follows the same pattern in a Java EE to Spring Boot migration. Spring Boot 3 requires Java 17 and Jakarta EE, so imports change from javax to jakarta packages [3].

Spring Boot 3 also makes Hibernate 6.1 the default, and Hibernate no longer supports switching back to the old ID generator mappings. Entity IDs, query behavior, and schema migration tool versions all need a test against real data.

Embedded SQL gets a search of its own. Queries built by string concatenation, dynamic SQL in stored procedures, and error handling that branches on vendor error codes are spread across the codebase and need a targeted scan.

How a Schema Change Spreads to Services, Reports, and Integrations

A schema change during modernization touches every consumer of the table. Services, reporting tools, ETL jobs, partner extracts, and ad hoc queries from other teams all read the same columns.

Three practices keep a schema change contained.

  • Consumer inventory per table. Query logs and code search identify every reader and writer before the change is designed.
  • Expand-and-contract changes. New columns and tables are added alongside the old ones, consumers move one by one, and the old structure is dropped after the last one moves.
  • Views as a compatibility layer. A view that presents the old shape keeps reports running during the transition.

A global credit-scoring firm migrated more than 1.5 million lines of legacy ETL code to Apache Spark and Airflow. More than 80% of the code transformation was automated. Functional parity testing compared legacy and Spark outputs for every module, and lineage reports confirmed no transformation logic was dropped. The program delivered a 55% TCO reduction, 60% faster time-to-market for new credit products, and zero data loss.

Map How Your Codebase Depends on Its Database

Dependency Map
Risk Heatmap
Modernization Plan (3-5 Days)
MedtronicClairULAB Systems+more

How to Sequence Code, Data, and Configuration Changes

The order of changes decides whether the application keeps working during the program. One sequence holds across the five example stacks.

  1. Baseline the legacy behavior. Capture inputs, outputs, and data state before any change, so every later step has a reference.
  2. Put data access behind one layer. Route database calls through a single layer in the legacy code, which gives later steps one place to change.
  3. Decide each database object. Apply keep, move, or rewrite to every procedure, trigger, and view.
  4. Expand the schema. Add the structures the new code needs next to the old ones.
  5. Move code in slices. Route one capability at a time to the new code, with its data path working on day one.
  6. Change configuration per slice. Connection strings, credentials, feature flags, and scheduled jobs move with the slice that uses them.
  7. Contract the schema. Drop old structures after the last consumer has moved.

During steps 4 to 6, both systems read and write the same data. Each slice needs one source of truth for its tables, either a shared database or synchronization with an agreed owner per table.

Synchronization adds latency and conflict rules, so slices that share heavily written tables move together.

Step 5 uses the routing layer of the strangler fig approach, and these data steps keep that routing working at the database.

How to Prove a Modernized System Behaves Like the Legacy System

Characterization tests capture what the legacy code does today and turn it into the acceptance contract for the new code. That contract covers the five layers where the application’s behavior shows up.

Five Layers to Compare When Validating Modernized Code

LayerWhat to CompareWhat It Catches
API responsesStatus codes, payloads, and headers for recorded requestsContract changes and serialization differences
Database stateRows inserted, updated, and deleted by each transactionMissing trigger logic, changed defaults, rounding on write
Batch and scheduled jobsOutput tables and record counts per runChanged ordering, date boundaries, skipped records
Generated filesExports, statements, and partner extracts at field levelEncoding, formatting, and column order
Downstream reportsTotals and row-level results for the same periodChanged joins, view definitions, and aggregation rules

Response comparison has a documented limit on writes. GitHub’s Scientist library runs old and new code side by side and compares results, and its documentation limits it to methods that don’t change data. For writes, it recommends writing to both systems and verifying at read time [4]. Database-state comparison covers that gap.

Where a modernized system's behavior shows up

How to Build Behavior Baselines Across a Whole Application

Baselines for a whole application come from recorded production behavior, in four steps.

  • Record. Capture production or production-like inputs at each entry point, with the resulting responses, data changes, and outputs.
  • Generate. Turn recordings into characterization tests per module. Gen AI drafts tests from the legacy code paths, and engineers approve them before they become the contract.
  • Compare. Replay the same inputs against the modernized system and diff each layer automatically.
  • Triage. Classify each mismatch as a defect, an approved behavior change, or noise such as timestamps and generated IDs.

Recorded production data gets masked before it leaves the production boundary. Approved behavior changes are logged with an owner, so the baseline remains the reference for everything else.

Validating Side Effects, Timing, and Production Integrations

Some behavior never appears in a test run and needs a separate plan.

  • Side effects. Emails, queue messages, and file drops are captured by sandbox endpoints and compared by content.
  • Timing. Scheduled jobs, time zone handling, and period-close rounding are tested on simulated dates, including month-end and year-end.
  • Production integrations. Partner feeds and payment or regulatory interfaces that fire in production get a parallel run through one full business cycle, with outputs compared before cutover.

Data-level checks follow the same techniques as any data migration validation.

Find Where Behavior Risk Concentrates

The $0 Modernization Assessment produces a risk and complexity heatmap for one representative codebase, showing where behavior risk concentrates before transformation starts.

Gen AI in Code Modernization and Where Engineers Review Its Output

Gen AI speeds up code modernization at three points, and each has a review step.

UseWhat Gen AI DoesReview Point
ComprehensionReads the whole codebase and reconstructs business rules, data flows, and dependenciesEngineers confirm reconstructed rules against known business behavior
Test generationDrafts characterization tests from legacy code pathsEngineers approve tests before they become the acceptance contract
TransformationConverts code, data access, and SQL to the target stackEngineers review each change as a diff, checked against the baseline

The acceptance contract comes from the legacy baseline. Tests generated from the modernized code add coverage, and a transformation is accepted when it matches the baseline across all five comparison layers.

Database logic is the review hotspot. Reviewers check transaction scope, null handling, and implicit type conversion in converted PL/SQL and T-SQL, where code can compile and return different results.

Code generated without full-codebase context misses callers in other repositories. Comprehension runs across the whole estate before transformation for that reason.

GitHub Copilot Modernization vs. Legacyleap for Legacy Code Modernization

IDE upgrade agents run framework and runtime upgrades from the developer’s environment. GitHub Copilot modernization is Microsoft’s agent for Java, .NET, and C++, and its documentation sets out what it covers.

For .NET, its published scenarios include .NET version upgrades and WebForms to Blazor, with data access skills for EF6 to EF Core and LINQ to SQL [5].

For Java, it converts SQL in application code during an Oracle to Azure PostgreSQL migration, and schema migration runs through a separate PostgreSQL extension [6]. Microsoft also describes upgrading many repositories at once [7].

The table compares that documented scope with Legacyleap’s on the work this guide covers.

Work TypeGitHub Copilot Modernization (Microsoft’s Documentation)Legacyleap
.NET and Java upgradesVersion and framework upgrades across many repositories.NET Framework to .NET Core LTS, and Java EE, EJB, and Struts to Spring Boot
ASP.NET WebFormsWebForms to BlazorWebForms to Blazor or React
Database logic and schemaApplication SQL conversion for Azure PostgreSQL, schema through a separate extension, no stored procedure skill in the .NET listPL/SQL logic moved into Java or .NET, with the data layer modernized alongside the UI and services
VB6, Delphi, and AngularJS sourcesNo listed scenarioVB6 to .NET or React, Delphi to .NET, AngularJS to Angular, React, or Vue
Migration targetAzureCloud-neutral targets, with the platform running inside the customer’s own infrastructure

The scope difference decides the fit.

  • GitHub Copilot modernization. Framework or runtime upgrades on .NET or Java applications headed for Azure, where database logic and schema stay in place.
  • Legacyleap. Programs that move database-resident logic, cross stacks, or deploy outside Azure, with behavior parity validated against the legacy baseline.

For a deeper dive, read more on Legacyleap vs Copilots.

What Drives Legacy Code Modernization Timeline and Cost

Timeline and cost estimates for legacy code modernization vary because the inputs vary. Two applications of the same size can differ by a factor of three when one keeps its business logic in stored procedures, carries circular dependencies, and has no tests.

DriverHow It Moves the EstimateHow to Measure It
Business logic in the databaseEach procedure and trigger needs a decision, a conversion, and its own testsCount of procedures, triggers, and views the application calls
Starting test coverageLow coverage adds baseline work before transformationCoverage report and count of untested entry points
Dependency depth and cyclesCircular dependencies block slicing and force larger cutoversDependency graph with a cycle count
Integrations and downstream consumersEach consumer adds a contract to preserve and validateInventory of APIs, feeds, reports, and jobs
Target-state ambiguityUndecided architecture causes rework mid-programCount of open target decisions at kickoff

These five counts set the estimate range for an application of any size. The same drivers shape any application modernization cost estimate.

Handoff Gaps Between Separate Code Modernization Tools

Separate tools for each phase leave gaps at the handoffs.

  • Context lost after assessment. The dependency map from assessment lives in a report, and the transformation tool rebuilds its own view of the code.
  • Tests that don’t match the transformation. Tests built by a separate tool target the old structure and break on each restructured module, which turns test maintenance into its own workstream.
  • Validation after the fact. Parity checks that start late find defects after the code has moved on, when each fix touches more of the system.

Each gap widens at the data layer, where assessment, transformation, and validation each see a different slice of the database.

How Legacyleap Runs Legacy Code Modernization Across the Full Lifecycle

Legacyleap is a Gen AI-powered legacy application modernization platform built on multi-agent orchestration. Its five agents run the Assess, Comprehend, Modernize, Validate, and Deploy lifecycle on one persistent model of the codebase, which carries context across the handoffs.

  • Assessment Agent. Produces a dependency map, technical debt report, risk indicators, modernization hotspots, and an effort and timeline estimate.
  • Documentation Agent. Reconstructs architecture, module boundaries, data flows, and integration maps, with business logic documented from the code itself.
  • Recommendation Agent. Recommends refactor, replace, or retain per module, with ordered migration phases and a target architecture.
  • Modernization Agent. Delivers converted code as diff-based pull requests, with migration gap reports and TODO markers on low-confidence areas. Roughly 70% of the modernization work is automated, and engineers review and direct the rest.
  • QA Agent. Generates unit, integration, regression, API, and functional tests, and produces behavior parity validation reports against the legacy baseline before cutover.

The platform modernizes the UI, services, and data layers in coordination. No agent merges, deploys, or executes code autonomously, and every change arrives as a diff for human review.

Full-codebase grounding covers polyglot estates across .NET, Java, frontend, and database code. All processing runs inside the organization’s own infrastructure.

The credit-scoring program above paired automated transformation with the same parity discipline, with auto-generated unit tests for every migrated module and zero data loss.

Next Steps for Planning a Legacy Code Modernization Program

Legacy code modernization holds up on three conditions. Database-resident logic has a decision before transformation, code and data move in a planned order, and the result matches a legacy baseline across every layer where behavior shows up.

The $0 Modernization Assessment produces a dependency and module map, a risk and complexity heatmap, and a modernization plan for one representative codebase. It takes 3 to 5 days, runs inside the client’s environment, and draws on more than 150 production-grade assessments.

A technical demo shows how the five agents take an approved plan through code changes and parity validation.

Plan the Code and Data Work Together

A $0 Modernization Assessment produces a dependency and module map, a risk and complexity heatmap, and a modernization plan for one representative codebase in 3 to 5 days, inside your own environment.

FAQ

Q1. What is the difference between legacy code modernization and code migration?

Code migration is one move between languages, frameworks, or platforms, and legacy code modernization is the program that sequences several of them. Each migration inside the program gets its own behavior baseline and cutover date.

Q2. How is legacy code modernization different from refactoring?

Refactoring restructures code with its behavior and platform unchanged. A refactor is one workstream inside a modernization program, which can also change the platform, the architecture, and the database.

Q3. Can AI modernize legacy code without human review?

No. Each change needs a named engineer accountable for it, and regulated systems need that approval recorded in the audit trail.

Q4. How long does legacy code modernization take?

A timeline comes from the delivery rate measured on a pilot module, applied across the estate’s driver counts. A pilot module with stored procedures and integrations keeps the measured rate realistic for harder modules.

Q5. Can legacy code be modernized without changing the database?

Yes, when the logic lives in application code and the database engine stays supported. Keeping the database fixed for the first release also separates code defects from data defects during validation.

Q6. What is a characterization test?

A characterization test records what existing code does today, including results that look wrong, and fails when that behavior changes. Michael Feathers named the technique in Working Effectively with Legacy Code, for code that has no specification.

References

[1] Forrester, “Forrester’s Technology & Security Predictions 2025” (press release)

[2] Microsoft Learn, “Port from EF6 to EF Core”

[3] Spring Boot, “Spring Boot 3.0 Migration Guide”

[4] GitHub, “Scientist”

[5] Microsoft Learn, “GitHub Copilot modernization scenarios and skills”

[6] Microsoft Learn, “Migrate from Oracle to PostgreSQL by using GitHub Copilot modernization”

[7] Microsoft Learn, “GitHub Copilot modernization agent overview”

Book a $0 Assessment

We will scan a portion of your legacy codebase and share documentation, architecture maps, dependency graphs, in 3-5 days.

Book a Time →
Share the Blog

Latest Blogs

MySQL to Aurora Migration: Paths, Upgrades, and App Changes

MySQL to Aurora Migration: Paths, Version Upgrades, and Application Readiness

Db2 to Postgres Migration: Schema, SQL PL, and App Code

Db2 to Postgres Migration: A Guide to Schema, SQL PL, and Application Code

SQL Server to PostgreSQL Migration: T-SQL, Apps, and Tools

SQL Server to PostgreSQL Migration: A Guide to T-SQL, Applications, and Tooling

Oracle to PostgreSQL Migration: Code, Data, and Validation

Oracle to PostgreSQL Migration: A Guide to Code, Data, and Validation

Legacy System Audit: What to Check Before Modernizing

Legacy System Audit: How to Assess Code and Data Before Modernization

Data Migration Validation, A Practitioner Guide

Data Migration Validation: A Practitioner’s Guide to Validating Data After Migration

Technical Demo

Book a Technical Demo

Explore how Legacyleap’s Gen AI agents analyze, refactor, and modernize your legacy applications, at unparalleled velocity.

Watch how Legacyleap’s Gen AI agents modernize legacy apps ~50-70% faster

Want an Application Modernization Cost Estimate?

Get a detailed and personalized cost estimate based on your unique application portfolio and business goals.