TL;DR
- Classify each application by friction and data coupling. Missing tests, missing documentation, and logic held in the database each add effort before transformation starts.
- Decide where database-resident logic lands before transformation starts. Each stored procedure, trigger, and view stays in the database, moves into service code, or gets rewritten for the target database.
- Sequence data changes with the code. Put data access behind one layer, change schemas in expand-and-contract steps, and move code in slices that each have a working data path.
- Prove behavior across five layers. Compare API responses, database state, batch outputs, generated files, and downstream reports against a legacy baseline before cutover.
- Estimate from drivers. Database logic, starting test coverage, dependency depth, integration count, and target-state clarity set the timeline.
What Legacy Code Modernization Involves
Legacy code modernization moves an application’s code to a current language, framework, or architecture, with the business running on it throughout. The code depends on a database that holds part of its logic, so the data layer moves in the same program. The work is complete when the modernized system produces the same results, data, and outputs as the legacy system.
Code migration moves code from one language, framework, or platform to another and keeps its behavior the same, as in Struts to Spring Boot or .NET Framework to .NET 8. A modernization program contains one or more code migrations, plus the architecture, data, and validation work around them.
Legacy is measured by friction. An application is legacy when the team can’t change it safely, because tests are missing, documentation is out of date, or the people who understood it have left.
Technical debt drives that friction. Forrester predicted that 75% of technology decision-makers would see their technical debt rise to a moderate or high level of severity by 2026 [1].
This guide covers distributed enterprise stacks, with examples drawn from five common legacy estates.
- .NET Framework and ASP.NET WebForms to modern .NET.
- Java EE and Struts to Spring Boot.
- AngularJS to Angular or React.
- VB6 and Delphi to .NET.
- PL/SQL- and T-SQL-heavy applications, where the database holds much of the logic.
How to Choose a Code Modernization Path for Each Application
Code modernization runs one decision per application. Business value, rate of change, and data coupling decide the path, and the data-layer column below shows how much database work comes with each choice.
| Path | What Changes | Data-Layer Impact | Fits When |
| Retain | Nothing for now | None | Low change rate on a supported runtime |
| Retire | The application is removed | Data is archived or merged elsewhere, and consumers are repointed | Functionality is duplicated or unused |
| Rehost | Infrastructure | Same database, new network path and connection settings | A data center exit sets the deadline |
| Replatform | Runtime or managed services, with minimal code change | Driver and connection changes, or a new database engine | A supported upgrade path exists |
| Refactor | Internal structure, same behavior | Data access layer cleaned up, schema largely unchanged | The domain model holds and tests can be built |
| Rearchitect | Boundaries, such as monolith to services | Schema split by service, database logic redistributed | Scale or team ownership needs new boundaries |
| Rebuild | Code rewritten on a new stack | Data model redesigned, full data migration | The stack has no upgrade path and requirements are recoverable |
| Replace | A commercial or SaaS product | Data moved into the vendor’s schema | The capability is a commodity |
Retire and replace remove code from the estate, so a portfolio review that starts with those two paths shrinks the scope before any transformation work is priced.
Four criteria decide between refactoring and rebuilding an application.
| Criterion | Points to Refactor | Points to Rebuild |
| Domain model | The model matches how the business works today | The business has outgrown the model |
| Upgrade path | The framework has a supported path forward | The stack has no upgrade path, as with VB6 |
| Testability | Characterization tests can be built around the code | The code is untestable in place, so tests wrap the legacy system’s inputs and outputs |
| Database coupling | Rules sit in application code | Rules sit in stored procedures and need redesign in the new data model |
The choice between refactoring and replatforming follows the same per-application logic.
Database Logic, Data Access, and Schema Dependencies in Legacy Code Modernization
Legacy applications keep part of their logic outside the application code. Three areas decide how much of a legacy code modernization program is data work.
- Business logic in the database. Stored procedures, triggers, and views that hold business rules.
- The data access layer. ORM mappings, embedded SQL, and transaction handling in the code.
- Schema dependencies. Every service, report, and integration that reads a table.

Business Logic in Stored Procedures, Triggers, and Views
Three kinds of database object carry business rules in older applications.
- Stored procedures. PL/SQL packages and T-SQL procedures hold pricing rules, eligibility checks, status transitions, and period-close calculations.
- Triggers. They enforce rules on every write, including writes from batch jobs and integrations the application team doesn’t own.
- Views. They encode the joins and filters that reports depend on.
Moving the application without deciding where that logic lands breaks behavior without a visible error. New code calls the same procedure with different parameters, or a trigger fires a second time once the service layer adds the same rule.
Each database object gets one of three decisions before transformation starts.
- Keep in the database. Set-based logic that runs close to the data, such as bulk updates and period-close calculations. It stays on the same engine and gets characterization tests of its own.
- Move into service code. Business rules that change often or need unit tests, such as pricing and eligibility. The procedure is retired once the service owns the rule.
- Rewrite for the target database. Logic that stays in the database through an engine change, such as PL/SQL packages moving to PL/pgSQL.
Engine changes add their own conversion work, covered for an Oracle to PostgreSQL migration, a SQL Server to PostgreSQL migration, and a Db2 to PostgreSQL migration.
Data Access Layer Changes in a Code Migration
Every code migration touches the layer between the code and the database. ORM mappings, raw SQL strings, connection and transaction handling, and vendor-specific error codes all change when the framework or the database changes.
.NET Framework to modern .NET shows the scale. Microsoft describes EF Core as a total rewrite of Entity Framework with no direct upgrade path [2]. EF6 Code First migrations history doesn’t carry over, so teams start migrations fresh. EDMX models need conversion first, since EF Core doesn’t support the format.
Applications built on raw ADO.NET add an ORM decision, covered in an ADO.NET to EF Core migration.
Java follows the same pattern in a Java EE to Spring Boot migration. Spring Boot 3 requires Java 17 and Jakarta EE, so imports change from javax to jakarta packages [3].
Spring Boot 3 also makes Hibernate 6.1 the default, and Hibernate no longer supports switching back to the old ID generator mappings. Entity IDs, query behavior, and schema migration tool versions all need a test against real data.
Embedded SQL gets a search of its own. Queries built by string concatenation, dynamic SQL in stored procedures, and error handling that branches on vendor error codes are spread across the codebase and need a targeted scan.
How a Schema Change Spreads to Services, Reports, and Integrations
A schema change during modernization touches every consumer of the table. Services, reporting tools, ETL jobs, partner extracts, and ad hoc queries from other teams all read the same columns.
Three practices keep a schema change contained.
- Consumer inventory per table. Query logs and code search identify every reader and writer before the change is designed.
- Expand-and-contract changes. New columns and tables are added alongside the old ones, consumers move one by one, and the old structure is dropped after the last one moves.
- Views as a compatibility layer. A view that presents the old shape keeps reports running during the transition.
A global credit-scoring firm migrated more than 1.5 million lines of legacy ETL code to Apache Spark and Airflow. More than 80% of the code transformation was automated. Functional parity testing compared legacy and Spark outputs for every module, and lineage reports confirmed no transformation logic was dropped. The program delivered a 55% TCO reduction, 60% faster time-to-market for new credit products, and zero data loss.
Map How Your Codebase Depends on Its Database


+moreHow to Sequence Code, Data, and Configuration Changes
The order of changes decides whether the application keeps working during the program. One sequence holds across the five example stacks.
- Baseline the legacy behavior. Capture inputs, outputs, and data state before any change, so every later step has a reference.
- Put data access behind one layer. Route database calls through a single layer in the legacy code, which gives later steps one place to change.
- Decide each database object. Apply keep, move, or rewrite to every procedure, trigger, and view.
- Expand the schema. Add the structures the new code needs next to the old ones.
- Move code in slices. Route one capability at a time to the new code, with its data path working on day one.
- Change configuration per slice. Connection strings, credentials, feature flags, and scheduled jobs move with the slice that uses them.
- Contract the schema. Drop old structures after the last consumer has moved.
During steps 4 to 6, both systems read and write the same data. Each slice needs one source of truth for its tables, either a shared database or synchronization with an agreed owner per table.
Synchronization adds latency and conflict rules, so slices that share heavily written tables move together.
Step 5 uses the routing layer of the strangler fig approach, and these data steps keep that routing working at the database.
How to Prove a Modernized System Behaves Like the Legacy System
Characterization tests capture what the legacy code does today and turn it into the acceptance contract for the new code. That contract covers the five layers where the application’s behavior shows up.
Five Layers to Compare When Validating Modernized Code
| Layer | What to Compare | What It Catches |
| API responses | Status codes, payloads, and headers for recorded requests | Contract changes and serialization differences |
| Database state | Rows inserted, updated, and deleted by each transaction | Missing trigger logic, changed defaults, rounding on write |
| Batch and scheduled jobs | Output tables and record counts per run | Changed ordering, date boundaries, skipped records |
| Generated files | Exports, statements, and partner extracts at field level | Encoding, formatting, and column order |
| Downstream reports | Totals and row-level results for the same period | Changed joins, view definitions, and aggregation rules |
Response comparison has a documented limit on writes. GitHub’s Scientist library runs old and new code side by side and compares results, and its documentation limits it to methods that don’t change data. For writes, it recommends writing to both systems and verifying at read time [4]. Database-state comparison covers that gap.

How to Build Behavior Baselines Across a Whole Application
Baselines for a whole application come from recorded production behavior, in four steps.
- Record. Capture production or production-like inputs at each entry point, with the resulting responses, data changes, and outputs.
- Generate. Turn recordings into characterization tests per module. Gen AI drafts tests from the legacy code paths, and engineers approve them before they become the contract.
- Compare. Replay the same inputs against the modernized system and diff each layer automatically.
- Triage. Classify each mismatch as a defect, an approved behavior change, or noise such as timestamps and generated IDs.
Recorded production data gets masked before it leaves the production boundary. Approved behavior changes are logged with an owner, so the baseline remains the reference for everything else.
Validating Side Effects, Timing, and Production Integrations
Some behavior never appears in a test run and needs a separate plan.
- Side effects. Emails, queue messages, and file drops are captured by sandbox endpoints and compared by content.
- Timing. Scheduled jobs, time zone handling, and period-close rounding are tested on simulated dates, including month-end and year-end.
- Production integrations. Partner feeds and payment or regulatory interfaces that fire in production get a parallel run through one full business cycle, with outputs compared before cutover.
Data-level checks follow the same techniques as any data migration validation.
Find Where Behavior Risk Concentrates
The $0 Modernization Assessment produces a risk and complexity heatmap for one representative codebase, showing where behavior risk concentrates before transformation starts.
Gen AI in Code Modernization and Where Engineers Review Its Output
Gen AI speeds up code modernization at three points, and each has a review step.
| Use | What Gen AI Does | Review Point |
| Comprehension | Reads the whole codebase and reconstructs business rules, data flows, and dependencies | Engineers confirm reconstructed rules against known business behavior |
| Test generation | Drafts characterization tests from legacy code paths | Engineers approve tests before they become the acceptance contract |
| Transformation | Converts code, data access, and SQL to the target stack | Engineers review each change as a diff, checked against the baseline |
The acceptance contract comes from the legacy baseline. Tests generated from the modernized code add coverage, and a transformation is accepted when it matches the baseline across all five comparison layers.
Database logic is the review hotspot. Reviewers check transaction scope, null handling, and implicit type conversion in converted PL/SQL and T-SQL, where code can compile and return different results.
Code generated without full-codebase context misses callers in other repositories. Comprehension runs across the whole estate before transformation for that reason.
GitHub Copilot Modernization vs. Legacyleap for Legacy Code Modernization
IDE upgrade agents run framework and runtime upgrades from the developer’s environment. GitHub Copilot modernization is Microsoft’s agent for Java, .NET, and C++, and its documentation sets out what it covers.
For .NET, its published scenarios include .NET version upgrades and WebForms to Blazor, with data access skills for EF6 to EF Core and LINQ to SQL [5].
For Java, it converts SQL in application code during an Oracle to Azure PostgreSQL migration, and schema migration runs through a separate PostgreSQL extension [6]. Microsoft also describes upgrading many repositories at once [7].
The table compares that documented scope with Legacyleap’s on the work this guide covers.
| Work Type | GitHub Copilot Modernization (Microsoft’s Documentation) | Legacyleap |
| .NET and Java upgrades | Version and framework upgrades across many repositories | .NET Framework to .NET Core LTS, and Java EE, EJB, and Struts to Spring Boot |
| ASP.NET WebForms | WebForms to Blazor | WebForms to Blazor or React |
| Database logic and schema | Application SQL conversion for Azure PostgreSQL, schema through a separate extension, no stored procedure skill in the .NET list | PL/SQL logic moved into Java or .NET, with the data layer modernized alongside the UI and services |
| VB6, Delphi, and AngularJS sources | No listed scenario | VB6 to .NET or React, Delphi to .NET, AngularJS to Angular, React, or Vue |
| Migration target | Azure | Cloud-neutral targets, with the platform running inside the customer’s own infrastructure |
The scope difference decides the fit.
- GitHub Copilot modernization. Framework or runtime upgrades on .NET or Java applications headed for Azure, where database logic and schema stay in place.
- Legacyleap. Programs that move database-resident logic, cross stacks, or deploy outside Azure, with behavior parity validated against the legacy baseline.
For a deeper dive, read more on Legacyleap vs Copilots.
What Drives Legacy Code Modernization Timeline and Cost
Timeline and cost estimates for legacy code modernization vary because the inputs vary. Two applications of the same size can differ by a factor of three when one keeps its business logic in stored procedures, carries circular dependencies, and has no tests.
| Driver | How It Moves the Estimate | How to Measure It |
| Business logic in the database | Each procedure and trigger needs a decision, a conversion, and its own tests | Count of procedures, triggers, and views the application calls |
| Starting test coverage | Low coverage adds baseline work before transformation | Coverage report and count of untested entry points |
| Dependency depth and cycles | Circular dependencies block slicing and force larger cutovers | Dependency graph with a cycle count |
| Integrations and downstream consumers | Each consumer adds a contract to preserve and validate | Inventory of APIs, feeds, reports, and jobs |
| Target-state ambiguity | Undecided architecture causes rework mid-program | Count of open target decisions at kickoff |
These five counts set the estimate range for an application of any size. The same drivers shape any application modernization cost estimate.
Handoff Gaps Between Separate Code Modernization Tools
Separate tools for each phase leave gaps at the handoffs.
- Context lost after assessment. The dependency map from assessment lives in a report, and the transformation tool rebuilds its own view of the code.
- Tests that don’t match the transformation. Tests built by a separate tool target the old structure and break on each restructured module, which turns test maintenance into its own workstream.
- Validation after the fact. Parity checks that start late find defects after the code has moved on, when each fix touches more of the system.
Each gap widens at the data layer, where assessment, transformation, and validation each see a different slice of the database.
How Legacyleap Runs Legacy Code Modernization Across the Full Lifecycle
Legacyleap is a Gen AI-powered legacy application modernization platform built on multi-agent orchestration. Its five agents run the Assess, Comprehend, Modernize, Validate, and Deploy lifecycle on one persistent model of the codebase, which carries context across the handoffs.
- Assessment Agent. Produces a dependency map, technical debt report, risk indicators, modernization hotspots, and an effort and timeline estimate.
- Documentation Agent. Reconstructs architecture, module boundaries, data flows, and integration maps, with business logic documented from the code itself.
- Recommendation Agent. Recommends refactor, replace, or retain per module, with ordered migration phases and a target architecture.
- Modernization Agent. Delivers converted code as diff-based pull requests, with migration gap reports and TODO markers on low-confidence areas. Roughly 70% of the modernization work is automated, and engineers review and direct the rest.
- QA Agent. Generates unit, integration, regression, API, and functional tests, and produces behavior parity validation reports against the legacy baseline before cutover.
The platform modernizes the UI, services, and data layers in coordination. No agent merges, deploys, or executes code autonomously, and every change arrives as a diff for human review.
Full-codebase grounding covers polyglot estates across .NET, Java, frontend, and database code. All processing runs inside the organization’s own infrastructure.
The credit-scoring program above paired automated transformation with the same parity discipline, with auto-generated unit tests for every migrated module and zero data loss.
Next Steps for Planning a Legacy Code Modernization Program
Legacy code modernization holds up on three conditions. Database-resident logic has a decision before transformation, code and data move in a planned order, and the result matches a legacy baseline across every layer where behavior shows up.
The $0 Modernization Assessment produces a dependency and module map, a risk and complexity heatmap, and a modernization plan for one representative codebase. It takes 3 to 5 days, runs inside the client’s environment, and draws on more than 150 production-grade assessments.
A technical demo shows how the five agents take an approved plan through code changes and parity validation.
Plan the Code and Data Work Together
A $0 Modernization Assessment produces a dependency and module map, a risk and complexity heatmap, and a modernization plan for one representative codebase in 3 to 5 days, inside your own environment.
FAQ
Code migration is one move between languages, frameworks, or platforms, and legacy code modernization is the program that sequences several of them. Each migration inside the program gets its own behavior baseline and cutover date.
Refactoring restructures code with its behavior and platform unchanged. A refactor is one workstream inside a modernization program, which can also change the platform, the architecture, and the database.
No. Each change needs a named engineer accountable for it, and regulated systems need that approval recorded in the audit trail.
A timeline comes from the delivery rate measured on a pilot module, applied across the estate’s driver counts. A pilot module with stored procedures and integrations keeps the measured rate realistic for harder modules.
Yes, when the logic lives in application code and the database engine stays supported. Keeping the database fixed for the first release also separates code defects from data defects during validation.
A characterization test records what existing code does today, including results that look wrong, and fails when that behavior changes. Michael Feathers named the technique in Working Effectively with Legacy Code, for code that has no specification.
References
[1] Forrester, “Forrester’s Technology & Security Predictions 2025” (press release)
[2] Microsoft Learn, “Port from EF6 to EF Core”
[3] Spring Boot, “Spring Boot 3.0 Migration Guide”
[5] Microsoft Learn, “GitHub Copilot modernization scenarios and skills”
[6] Microsoft Learn, “Migrate from Oracle to PostgreSQL by using GitHub Copilot modernization”
[7] Microsoft Learn, “GitHub Copilot modernization agent overview”







