← Back to list
SW Engineering & Management
#3R#역공학#재공학#재구조화#레거시#133회
Last updated · 2026-09-30

Software Maintenance 3R (Reverse Engineering · Restructuring · Reengineering)

1. Overview

A. Definition

A collective term (3R) for the three representative techniques of analyzing (reverse engineering), improving (restructuring), and rebuilding (reengineering) an existing system, in order to raise the maintainability of aging, complex software and reduce its total cost of ownership (TCO).

Rather than three independent techniques, 3R comprises activities placed on a single continuum (spectrum) running from "analysis → improvement → rebuilding." If reverse engineering is the stage of "understanding what this system is and how it was built," restructuring is the stage of "tidying up what was understood into something better while keeping the function unchanged," and reengineering is the stage that includes both of the former and "builds it anew." Therefore, among the three techniques reengineering is the most comprehensive, and one must first grasp that it embraces reverse engineering and restructuring as its own sub-processes so that the conceptual layering does not become muddled.

The reason these concepts are treated together is that, when working on legacy systems in practice, these three activities almost always appear together. An undocumented system must first have its structure recovered through reverse engineering; the recovered structure usually has many places to fix, so restructuring follows; and if the platform must be changed as well, it expands into reengineering. In other words, 3R can be seen as a spectrum of intensity aimed at the single goal of legacy modernization — reverse engineering is the lightest intervention, restructuring the middle, and reengineering the heaviest.

B. Background and Necessity

A long-operated legacy system reaches a state where no one in the organization fully understands its structure anymore, as staff turn over repeatedly and design documents are lost. On top of this, when feature additions and emergency fixes accumulate over a long period, the code's coupling rises and its cohesion falls, and so-called technical debt compounds like interest. In this state, even a single line of change triggers an unexpected ripple effect, and the cost of verifying regressions per change increases exponentially. In fact, the point that a substantial share of a software's total lifecycle cost (commonly cited as more than half) arises in the post-development maintenance phase illustrates well the weight of this problem.

Yet if one sweeps away the entire system and goes for a new development (rewrite), the business rules (domain knowledge) accumulated implicitly only inside the code over many years are lost, and the risk and cost of running the old and new systems in parallel become very large. This is why large-scale rewrite projects often fail or greatly overshoot their schedules. Between these two extremes (neglect vs. full rewrite), 3R offers a compromise that reuses existing assets as much as possible while lowering risk and cost to a controllable level. Amid today's trend of Legacy Modernization, which moves on-premises legacy to the cloud and microservices (MSA), 3R is drawing renewed attention as its core means of execution.

C. Characteristics

The first characteristic of 3R is asset reusability. All three techniques make the most of existing code, data, and business rules rather than building from scratch, so improvement is achieved while preserving proven domain knowledge. The second characteristic is that they presuppose the preservation of functional equivalence: reverse engineering and restructuring do not change the visible function at all, and reengineering, too, adds enhancements only after verifying that existing functions are preserved. The third characteristic is that they form a continuous spectrum — reverse engineering, restructuring, and reengineering are stages of steadily increasing intervention intensity, and one selects or combines only the intensity needed to match the nature of the problem.

2. The Overall Structure of 3R and Criteria for Distinction

To use 3R accurately, it is clearest to distinguish the three techniques by how they move across the level of abstraction. Software has an abstraction ladder of requirements → design/specification → implementation (code), and each technique moves in a different direction on this ladder.

flowchart LR
  L["Legacy SW(code, artifacts)"] --> RE["Reverse Engineering(RE)"]
  RE --> RS["Restructuring(RS)"]
  RS --> RN["Reengineering(RN)"]
  RN --> N["Modernized SW"]
  RE -. "recover upper spec" .-> SPEC["design, spec"]
  SPEC -. "re-implement via forward eng." .-> RN

Reverse engineering is a bottom-up activity that climbs back up from implementation (code) to the higher concepts of design and specification. Its purpose is to revive the data structures, control flows, and business rules contained within by reading the code and turning them into diagrams or specifications, and it does not change the externally visible behavior of the system at all. It is the "starting point of understanding" that must be passed through when dealing with undocumented legacy, and it determines the accuracy of every improvement that follows.

Restructuring is an activity that improves only the representation within the same layer without moving the level of abstraction. For example, it stays in the code layer to split spaghetti code into modules, or stays in the data layer to tidy up a schema whose normalization has broken. Its core is to raise internal quality while keeping the visible function identical, so its effect appears not as immediate features but as a reduction in the cost of later changes.

Reengineering is a "bottom-up → top-down" round-trip activity that recovers the upper specification through reverse engineering, improves that specification, and comes back down to implementation via forward engineering. It is the only one of the three techniques that can change even the visible function, quality, and platform, and its scope and risk are correspondingly large. Reengineering is therefore the heaviest intervention, one that must be preceded by a business justification for "why rebuild now."

Technique Main purpose Change in visible function Direction of abstraction movement Representative deliverable
Reverse engineering Extract design/spec (understand) None Upward (code→design) Recovered design/data model
Restructuring Improve structure/readability/complexity None Same level Improved code/schema
Reengineering Rebuild/quality/platform upgrade Possible Upward→downward (round-trip) Rebuilt system

As the table shows, the two axes that divide the three techniques are "does it change the visible function" and "where does it move on the abstraction ladder." Answering these two questions naturally decides which technique to use in a given situation.

3. An In-Depth Understanding of Each Technique

A. Reverse Engineering

Reverse engineering is the activity of extracting and recovering design and specifications from already-built artifacts such as source code, executable binaries, databases, and screens. It brings into software the concept, originally from manufacturing, of disassembling a finished product to figure out its design, and its core lies not in "building" but in "understanding." Reverse engineering therefore never changes the system's behavior and concentrates solely on reviving knowledge.

The targets of reverse engineering are broadly divided into a code view and a data view. From the code view, it recovers control-flow graphs, call relationships, and class diagrams; from the data view, it revives a logical data model (ERD) from the physical schema. In practice, static analysis (analyzing structure without running the code) and dynamic analysis (observing actual flow via runtime logs and profiling) are used together to raise accuracy. For example, when dealing with a 20-year-old COBOL core system, static analysis alone makes it hard to know which branches are actually live, so "dead code" is identified through dynamic analysis based on operational logs.

The reason reverse engineering matters is that the success of legacy modernization hinges on "how accurately it was understood." If understanding is poor, the restructuring and reengineering that follow are built on wrong premises and damage the original function. Recently, large language model (LLM)-based code summarization and explanation tools have greatly accelerated the speed of understanding in reverse engineering, but automated recovery results can contain errors, so human verification must always follow.

B. Restructuring

Restructuring is the activity of improving only the internal code, structure, and representation while keeping the externally visible function and behavior unchanged. Its purpose is to raise readability and modularity and lower cyclomatic complexity and duplication, thereby making later changes easier and safer. The value of restructuring is not "delivering a new feature right now" but "lowering the cost of future change," so its effect appears not as immediate revenue but as maintenance productivity.

Typical restructuring tasks include splitting long functions into meaningful units, extracting duplicated logic into shared modules, flattening deeply nested conditionals, and turning magic numbers into constants. On the data side, normalizing tangled denormalized tables or restoring referential integrity constraints falls here. The important premise is that the visible behavior must be completely identical before and after restructuring, and the safeguard that guarantees this is regression testing. In legacy with insufficient tests, the standard approach is to first secure characterization tests to "pin" the current behavior before touching anything.

Restructuring is often confused with refactoring, but the two are on different layers. Restructuring is a broad concept spanning code, data, and architecture, while refactoring is the practical technique among these of applying small units repeatedly at the source-code level. For example, in a large banking system "reorganizing the payment module into a layered structure" is restructuring, and splitting individual functions and renaming them in that process is refactoring.

C. Reengineering

Reengineering is the most comprehensive activity, combining reverse engineering + improvement + forward engineering into one to rebuild a system. It first recovers the design and business rules of the existing system through reverse engineering, then removes defects, duplication, and outdated design from the recovered specification to improve it, and re-implements it via forward engineering to fit a new platform, language, and architecture based on the improved specification. The decisive difference from restructuring is that, in this process, the visible function may improve or the platform may change entirely.

The strength of reengineering is that it can fundamentally improve the system while preserving accumulated domain knowledge. For example, a project moving a mainframe COBOL core system to a Java-based web system is a classic reengineering effort that does not re-imagine screens and data from scratch but re-implements existing business rules on the new platform after securing them through reverse engineering. This way, only the technology stack is modernized while the proven knowledge of "what must be calculated" is retained.

However, since reengineering has the largest scope and risk of the three techniques, clear goal-setting and a phased migration strategy for what to change and how far are essential. A commonly used approach is not a big-bang replacement of the whole system at once, but the Strangler Fig pattern, which gradually migrates functionality unit by unit to the new system while placing a relay layer in front of the old system. This approach does not stop the service during migration and can roll back only the affected function if a problem arises, greatly lowering risk.

4. The Reengineering Process and Related Concepts

How reengineering weaves the three techniques into a single flow can be seen in the following process diagram.

flowchart TD
  A["Analyze / reverse-eng (recover structure, rules)"] --> B["Improve / restructure (remove defects, duplication)"]
  B --> C["Transform / forward-eng (re-implement on new platform)"]
  C --> D["Test / transition (regression verify, parallel run)"]
  D --> E["Roll into ops, stabilize"]
  D -. "functional mismatch found" .-> A

Reengineering first analyzes and reverse-engineers to recover the structure and business rules of the existing system, then improves and restructures to remove design defects and duplication, then transforms and forward-engineers to re-implement it to fit the new platform and structure, and finally tests and transitions to verify that existing functions are preserved and reflect it into operation. The most important control point in this flow is the final testing stage. It must be confirmed via regression testing that the original function operates identically even after rebuilding, and it has an iterative structure that returns to the analysis stage when a functional mismatch is found. In practice, this verification is carried out at scale through a parallel run, which feeds the same input into the old and new systems and compares the outputs, and the old system is retired only when the calculation results match down to the smallest unit.

3R is often confused with adjacent concepts, so it is worth organizing the relationships. In particular, knowing where forward engineering, migration, and refactoring sit within 3R makes the concepts clear.

Concept Description Relationship to 3R
Forward Engineering Forward development of spec → design → implementation The final process of reengineering
Migration Moving platform/language/DB to a different environment A form/part of reengineering
Refactoring Improving internal structure while keeping external behavior Practicing restructuring at the code level

Refactoring is the practical technique of applying the concept of restructuring repeatedly in small units at the source-code level, and migration can be seen as a special case of changing the target environment (platform/language/DB) during the reengineering process. Thus, saying "we migrated" usually means one performed part of reengineering, and saying "we refactored" means one practiced restructuring locally.

Positioning adjacent concepts on the coordinates of 3R this way can reduce confusion in practical discussion. In the field "reengineering" and "rewrite" are often used interchangeably, but a rewrite discards existing assets and builds from scratch, so its orientation differs from 3R, which presupposes asset reuse. A professional engineer should use these terms with precise distinction in proposals and design documents, aligning stakeholders' expectations about the project's scope, risk, and cost from the start.

5. Criteria for Selecting a Technique and Application Cases

Which of the three techniques to apply must be judged by the nature and goal of the problem, not by trend. If the problem is "no one knows the structure," reverse engineering comes first; if it is "we know the structure but are afraid to touch it," restructuring; and if it is "the platform itself has reached end of life," reengineering is the answer. In other words, jumping straight into reengineering without a diagnosis is like performing major surgery without knowing the cause of the illness.

To aid the judgment, look at three signals. The first is change frequency — the more frequently a module changes, the faster the return on investment in restructuring/reengineering. The second is defect density — if failures are concentrated in a particular module, that part is the priority target for improvement. The third is staff comprehension — an area no one understands must first have its knowledge recovered through reverse engineering. The module in which all three signals are high sits at the very front of the 3R investment priority.

As a concrete example, suppose a distributor's settlement batch is delayed by several hours at every month-end close and the person in charge has resigned, so no one can explain the logic. If one starts a rewrite right away here, there is a great risk of losing unverified settlement rules. The correct order is to first recover the batch's calculation rules and data flow through reverse engineering and leave them as a specification, pin the current output values with characterization tests, then restructure the bottleneck of repeated query logic, and if necessary replace the batch engine itself through reengineering. Following these steps this way, one can safely remove the root cause behind the outwardly visible symptom of "several hours of month-end delay."

As another case, a system whose specific language/runtime is nearing End of Support has its security patches cut off, so reengineering is effectively forced. In this case the primary goal becomes "reproducing equivalent functionality on a supported platform" rather than functional improvement, and here too the specification secured through reverse engineering and regression tests serve as the safety net for the transition. Ultimately the three techniques are used alone or in combination depending on the situation, but in common they must keep the order of "understand → verify → improve" to avoid failure.

6. In Depth: Linkage with Cloud Modernization Strategy and Recent Trends

3R has recently met cloud-transition discourse and expanded into a broader execution strategy. The 6R/7R framework widely cited in cloud migration (Rehost·Replatform·Refactor·Rearchitect·Rebuild·Replace, plus the extension adding Retain·Retire) is effectively a spectrum of "how deeply to touch the legacy," and of these, Refactor·Rearchitect·Rebuild directly abut 3R's restructuring and reengineering. For example, Rehost (lift-and-shift), which moves only the server to the cloud, uses almost no 3R, but Rearchitect, which reorganizes the application structure to fit containers and MSA, mobilizes reverse engineering, restructuring, and reengineering all together.

As a practical application case, the construction of next-generation systems at financial and public institutions at home and abroad is mostly a reengineering project of mainframe core systems. While moving COBOL assets on the scale of millions of lines to Java and the cloud, an approach in which an automated conversion tool generates first-cut code and humans then verify and correct the business rules has become standardized. Here, reconciliation of whether the old and new systems' calculation results match down to the won (₩) during the parallel-run period becomes the gateway to a successful transition.

The most notable recent trend is the spread of AI-based code analysis and automated conversion. LLM-based tools summarize the intent of unfamiliar code in the reverse-engineering stage, suggest refactoring candidates in the restructuring stage, and produce drafts that move legacy languages to modern languages in the conversion stage. This clearly reduces the time it takes for humans to read and understand code. However, automated conversion results can contain subtle semantic errors or hallucinations, so quality can be guaranteed only by always going through test-based verification and final human confirmation. From a professional engineer's perspective, a balance is required that positions AI as "an accelerator of understanding and draft generation" while still placing final responsibility for accuracy on the verification system.

7. Considerations and Implications

From a professional engineer's perspective, the success of 3R ultimately depends on the judgment of "what, why, and how far to touch." The following four points should be considered strategically.

  • Prioritizing targets (trade-off): Not all legacy can be reengineered. One must evaluate the portfolio along the axes of maintenance cost, business importance, technical-debt level, and change frequency, and touch first the systems with the highest cost-effectiveness. For a stable system that rarely changes, Retain may actually be reasonable.
  • A system to guarantee functional preservation: The greatest risk of 3R is losing unverified business rules during the rebuild. One must therefore build a safety net in advance that pins current behavior with characterization tests and proves functional equivalence through regression testing, parallel runs, and result reconciliation.
  • A phased migration strategy: Because a big-bang transition carries great risk, it is desirable to choose a structure that migrates gradually by functional unit and can roll back on problems, like the Strangler Fig pattern. This achieves service continuity and risk control at the same time.
  • Linkage with modernization strategy and AI use: 3R is the execution engine of cloud/MSA transition (6R/7R), and AI-based analysis/conversion tools can greatly raise the efficiency of reverse engineering and conversion. However, automation results must always presuppose verification, and the participation of staff who know the domain must be maintained to secure the quality and sustainability of modernization.

In one line: 3R is a legacy-improvement spectrum running from reverse engineering (design extraction) · restructuring (structural improvement) · reengineering (rebuilding), distinguished by the direction of abstraction movement and whether the function changes, and—premised on target prioritization, regression-test-based functional preservation, and phased transition—it becomes the core means of execution for cloud/MSA modernization (6R/7R) and AI automation.