Prepare for JCL interview questions grouped by experience level.
JCL Interview Question & Answers
0-2 Years
JCL stands for Job Control Language, and it is used on IBM mainframe systems to tell the operating system which programs to run, what datasets they need, and how much resource to allocate to them. It acts as the instruction set between a batch job and z/OS, since programs like COBOL executables cannot start themselves or find their own input and output files.
The three core statement types are JOB, EXEC, and DD. The JOB statement identifies the job and its accounting information, EXEC specifies which program or procedure to run, and DD (Data Definition) describes the datasets the program will read from or write to.
The JOB statement is always the first statement in a job and marks the start of a new unit of work for the system. It carries the job name, accounting information, programmer name, and parameters like CLASS and MSGCLASS that control how the job is scheduled and where its output goes.
The EXEC statement names the program or cataloged procedure that the step will run. A single job can have multiple EXEC statements, and each one defines a separate job step that executes in sequence.
A DD statement defines a dataset that a job step needs, either for input, output, or as a temporary work file. It specifies things like the dataset name, disposition, space allocation, and record format so the system knows exactly how to locate or create the file.
A job is the complete unit of work submitted to the system, identified by one JOB statement, while a job step is a single program execution within that job, identified by one EXEC statement. A job can contain several steps that run one after another, each potentially depending on the outcome of the step before it.
DISP controls the disposition of a dataset, meaning its status when the step starts and what should happen to it when the step ends. It is usually written as three sub-parameters like DISP=(NEW,CATLG,DELETE), covering the current status plus the normal and abnormal termination actions.
The common status values are NEW for a dataset being created in this step, OLD for exclusive access to an existing dataset, SHR for shared read access to an existing dataset, and MOD for appending to an existing dataset or creating it if it does not exist.
A PROC is a cataloged or in-stream procedure, which is a reusable block of JCL steps that can be invoked from an EXEC statement instead of writing the same steps repeatedly. It saves time and reduces errors because common processing logic is written once and referenced by name across many jobs.
A cataloged procedure is stored permanently in a procedure library like PROCLIB and referenced by name from any job, while an in-stream procedure is defined directly inside the job stream between PROC and PEND statements and only exists for that job run. In-stream procedures are handy for testing changes before promoting them to a cataloged version.
A condition code, also called a return code, is a numeric value a program sets when it finishes to indicate how it went, with 0 typically meaning success and higher numbers indicating warnings or errors. Later steps in the job can check this return code with the COND or IF/THEN logic to decide whether to run.
SYSOUT tells the system to route the dataset's output to a system output class rather than a permanent dataset, most commonly used for print listings, job logs, and reports. The class letter after SYSOUT= determines where and how that output gets processed, such as SYSOUT=* which follows the job's default message class.
SPACE tells the system how much storage to allocate for a new dataset, specified in units like tracks, cylinders, or blocks along with primary and secondary allocation amounts. Getting SPACE wrong is a common beginner mistake, since too small an allocation causes a job to abend with a space error partway through.
GDG stands for Generation Data Group, a way of managing a series of related datasets like daily transaction files under one base name with numbered generations. Referring to a dataset as +1 for the newest generation or 0 for the current one lets jobs process history without hardcoding date-stamped names.
UNIT=SYSDA tells the system to allocate the dataset on any available direct access storage device group defined as SYSDA at the installation, rather than a specific volume. It is one of the most common unit names for temporary work datasets and new datasets that do not need to sit on a particular disk pack.
A temporary dataset exists only for the duration of the job, usually named with a double ampersand like &&TEMP, and gets deleted automatically once the job ends. A permanent dataset has a fully qualified catalog name and persists after the job finishes so other jobs or users can access it later.
RECFM specifies the record format of a dataset, describing how records are structured on disk. Common values include FB for fixed-block records of a set length, VB for variable-block records that can differ in length, and F for fixed unblocked records.
LRECL stands for logical record length, and it tells the system the length in bytes of each record in the dataset. It has to match what the reading or writing program expects, since a mismatch between LRECL and the program's actual record layout leads to truncated or garbled data.
MSGCLASS sets the output class for the job's own system messages and JES log, separate from any SYSOUT the program itself produces. Operators commonly route MSGCLASS to a hold class so job output can be reviewed on screen instead of printing automatically.
If a program tries to open a dataset it expects but no matching DD statement was coded, the job typically abends with an error such as an open error or a JCL error depending on when the mismatch is caught. This is one of the most common beginner mistakes, since the ddname in the JCL has to exactly match what the program references internally.
A comment statement starts with //* and lets a developer add explanatory notes directly in the JCL without affecting execution. Comments are useful for documenting why a step exists, what a dataset is used for, or leaving a note for whoever maintains the job later.
CATLG tells the system to keep the dataset and add an entry for it in the system catalog so it can be referenced by name later without specifying its volume, while KEEP retains the dataset physically but does not catalog it, meaning later jobs would need to know its exact volume serial to find it. Most permanent datasets use CATLG for convenience.
PARM passes a short string of parameters directly to the program at execution time, which the program can read to change its behavior without needing a separate input file. It is commonly used for things like passing a run date or a processing mode flag to a COBOL program.
JES2 and JES3 are job entry subsystems that manage the flow of jobs through the mainframe, handling job submission, scheduling, spooling of output, and print management. JCL is the language used to describe jobs to whichever job entry subsystem the installation runs, and most day to day JCL coding looks identical regardless of which one is underneath.
Abend is short for abnormal end, referring to a job or step terminating unexpectedly due to an error rather than completing normally. Abends are usually reported with a system completion code like S0C7 for a data exception, which points the programmer toward the type of problem that occurred.
The step name is the label given to a job step, written before EXEC on the statement, and it lets later steps or condition code checks refer back to that specific step's outcome. Referencing a prior step's dataset in a later DD statement, such as with a backward reference, also relies on the step name.
DSN stands for data set name, and it identifies the fully qualified name of the dataset being referenced or created, such as PROD.PAYROLL.MASTER. Getting the naming convention right matters a lot on shared mainframe systems since naming standards usually map to ownership and retention rules.
REGION specifies the amount of virtual storage a job step is allowed to use, expressed in kilobytes or megabytes. If a program needs more memory than the REGION value allows, it will abend with a storage-related error, so REGION often needs adjusting for memory-intensive programs like sort utilities.
IEFBR14 is a tiny IBM utility program that does nothing except return a zero condition code immediately, and it is commonly used in a JCL step purely to allocate, delete, or rename a dataset without actually running real processing logic. It is a convenient way to create or clean up datasets as a standalone step.
Dataset organization describes how records are physically arranged in a file, and DSORG is the JCL keyword that can specify it, with common values like PS for physical sequential and PO for partitioned. Most modern JCL infers organization from context or the DCB subparameters rather than coding DSORG explicitly.
VOL specifies the volume serial number of the disk or tape where a dataset resides or should be placed, letting the system locate uncataloged datasets or direct new allocations to a specific device. It is less commonly needed for cataloged datasets since the catalog already tracks volume information.
SYSIN is a DD name conventionally used to feed input data or control statements into a program, often through an in-stream data set marked with an asterisk, while SYSOUT routes a program's printed or logged output to a system output class. The two names are just convention, but nearly every mainframe shop follows them consistently.
An in-stream dataset is data supplied directly within the job stream itself, starting after a DD * statement and ending with a delimiter line, typically /*. It is commonly used to pass small amounts of control card input or test data without needing a separate permanent file.
A JCL error means the job control language itself has a syntax problem, a missing required parameter, or references something the system cannot resolve, and it is caught before the program ever actually starts executing. An abend, by contrast, happens after the program has started running and something goes wrong during actual processing.
A null statement, written as just two slashes with nothing after them, marks the logical end of a job's JCL stream. It is optional in most environments since the next JOB statement or end of input naturally signals the boundary, but some shops still include it as a convention.
TYPRUN changes how a job is processed instead of running it normally, with common values like HOLD to keep it in the queue until released, SCAN to check the JCL for syntax errors without executing, and JCLHOLD to hold it for JCL checking. It is a handy way to validate new JCL without risking an actual run against production data.
3-6 Years
I would start by pulling the abend dump or SDUMP output and looking at the offset where the failure occurred, since S0C7 is a data exception almost always caused by invalid numeric data in a COMP-3 or display numeric field. From there I would trace back to the input dataset to find which record has non-numeric characters in a field the program treats as packed decimal, since bad data upstream is the most common root cause.
I would use a PARM parameter on the EXEC statement to pass the date string, and have the program's linkage section pick it up so the same JCL works day after day without editing. For more complex parameter sets I would instead point a DD statement at a small parameter dataset the program reads at startup, since that scales better than PARM's length limits.
I would use COND or, in more recent JCL, an IF/THEN/ELSE/ENDIF block that checks the return code of the earlier step and only executes when that condition is met. IF/THEN logic is generally preferred over COND today because its syntax is much easier to read and maintain, especially once several conditional branches are involved.
I would raise both the primary and secondary quantities in the SPACE parameter, giving enough secondary extents that the dataset can grow without repeatedly hitting a B37 abend. I would also check whether switching from tracks to cylinders makes more sense for a large dataset, since cylinder allocation reduces the number of extents needed and improves I/O efficiency.
I would check the SPACE parameter on the DD statement that failed and compare it against how much data the step is actually writing, since B37 means the dataset ran out of allocated space and could not get a new extent. If the dataset already has the maximum number of extents, I would consider switching to cylinder allocation or splitting the output across multiple datasets.
I would give the DD statement in the creating step a step name, then in the later step reference it with DSN=*.stepname.ddname as a backward reference. This avoids retyping the full dataset attributes and guarantees the later step points at exactly the dataset the earlier step produced.
I would identify the steps that are genuinely reused across multiple jobs and move those into a cataloged procedure with symbolic parameters for anything that varies, like dataset qualifiers or a run indicator. Then each calling job's JCL becomes a short EXEC PROC=xxx statement with overrides, which cuts down maintenance since a logic change only needs to happen in one place.
I would define symbolics like &INFILE or &RUNDATE inside the proc's DD and EXEC statements, then supply actual values either as defaults on the PROC statement or as overrides on the calling EXEC statement. This lets one proc serve many different jobs since the calling JCL only needs to pass in the values that differ for that run.
I would check whether it is an enqueue conflict, meaning another job or user already has exclusive access to a dataset this job needs, which is common when DISP=OLD is requested on a dataset someone else is updating. I would also check job class and initiator availability, since a job can sit in the input queue simply because no initiator is free to run its assigned class.
I would use the RESTART parameter on the JOB statement to specify the step name to resume from, avoiding re-running earlier steps that already completed successfully and possibly already updated a file. Before restarting I would also check whether any dataset from the failed step needs to be cleaned up or reset to avoid duplicate processing.
I would use IDCAMS REPRO or IEBGENER depending on the dataset type, choosing IDCAMS for VSAM files and IEBGENER for simple sequential copies, and tune the BLKSIZE and buffer settings for large volumes. For very large datasets I would also consider a sort utility's copy function since some sort products handle bulk copies faster than IEBGENER.
I would write an IF/THEN block that evaluates a compound condition, such as IF STEP1.RC > 4 OR STEP2.RC > 0 THEN, and nest ELSE branches for the alternate paths. This is far more readable than chaining multiple COND parameters, especially once more than two conditions are involved.
I would code two DD statements referencing the same GDG base, one with (0) for the current generation and one with (-1) for the prior generation, letting the program compare the two in a single step. This pattern is common for delta processing where a job needs to identify what changed between yesterday's and today's file.
I would look at increasing BLKSIZE to reduce the number of physical I/O operations, check whether the job can use a larger REGION to avoid page faulting, and consider whether the processing logic itself could be handled by a sort utility with an embedded exit instead of a custom program. Parallelizing across job steps or splitting the file for concurrent processing is another option if the downstream logic allows it.
I would check for a shared dataset being updated concurrently by another job, since inconsistent output timing often points to a data dependency or timing race rather than a JCL problem itself. I would also review whether the job depends on a GDG generation that a scheduled job elsewhere creates, since running before that generation exists changes what data gets picked up.
I would rely on the job scheduler, such as CA-7 or Control-M, to trigger a post-processing action based on the job's return code rather than trying to build alerting into raw JCL itself, since JCL has no native email capability. Within the JCL I would make sure return codes are set accurately and consistently so the scheduler's condition logic can act on them reliably.
I would first inventory current naming patterns and dataset qualifiers to understand what conventions already exist informally, then propose a documented standard covering job names, step names, and dataset high-level qualifiers. I would roll it out gradually through new development and scheduled migrations rather than a risky mass rename of production jobs.
I would generate the JCL dynamically from a driver script or use a scheduler's variable substitution to inject the correct dataset name before submission, since raw JCL itself has no built-in day-of-week logic. Alternatively I would use a GDG with a consistent naming pattern so the same relative generation reference always resolves to the right day's file.
I would run the changed job in a test region against a representative copy of production data, compare record counts and key totals against a known baseline, and review the job log for unexpected return codes or messages. I would also check that any changed SPACE or BLKSIZE values did not silently truncate or reject records during the test run.
I would investigate adding an explicit dependency in the scheduler so this job only triggers after the upstream job's completion event, rather than relying on a fixed time window that can drift. If DISP conflicts are the direct cause, I would also confirm both jobs are not both requesting exclusive access to the same dataset unnecessarily.
I would change the UNIT parameter from a tape unit name to a disk group like SYSDA and add or adjust the SPACE parameter, since tape allocations do not use SPACE the same way disk does. I would also review whether stacking multiple temporary datasets under one unit affinity still makes sense once the medium changes.
I would code multiple DD statements with the same ddname stacked one after another, which JCL treats as a logical concatenation the program reads through sequentially. I would make sure the record formats and lengths across the concatenated datasets are compatible, since JCL does not enforce that they match.
I would replace the plain DD dataset references with the appropriate AMP and access method parameters for VSAM, since VSAM datasets need different subparameters than physical sequential ones. I would also confirm whether the program itself needs code changes for VSAM I/O, since the JCL update alone is not sufficient if the access method in the program does not match.
I would check whether the requested change is genuinely needed by all callers or just one, since adding a step unconditionally into a widely shared proc risks unintended side effects for jobs that do not need it. If the need is specific to one caller, I would make the new step conditional through a symbolic parameter so existing callers default to their current behavior and only opt in explicitly.
6-8 Years
I would build the dependency graph explicitly in the job scheduler using predecessor and successor relationships rather than relying on fixed run times, since interdependent batch cycles drift and fixed timing creates fragile races. I would group jobs into logical streams by business function, add checkpoint jobs that validate data readiness between streams, and design restart points so a failure in one stream does not force a full cycle rerun.
I would profile the critical path through the job dependency graph to find which chain of jobs actually determines the total elapsed time, since optimizing jobs off the critical path does not help the overall SLA. From there I would look at candidates for parallelization, larger REGION or BLKSIZE tuning on I/O-heavy steps, and whether any jobs are serialized unnecessarily due to overly broad dataset locking.
I would maintain versioned procedure libraries with a promotion process through test, QA, and production regions, and stage new procs under a temporary name before cutting jobs over, so a rollback is just repointing the EXEC statement. I would also require a change window that avoids mid-cycle promotion, since swapping a proc definition while dependent jobs are actively running risks inconsistent behavior.
I would structure the scheduler's dependencies so that only jobs genuinely dependent on the failed job's output are held, while unrelated streams continue processing in parallel. Within the JCL itself I would use IF/THEN logic and condition codes consistently across all jobs so the scheduler can reliably distinguish a hard failure from a warning that should not halt anything.
I would weigh the reduced initiator and scheduling overhead of fewer larger jobs against the loss of restart granularity, since a failure partway through a large consolidated job means more reprocessing than a failure in one small job would. I would generally keep jobs separated at natural business checkpoints and only consolidate steps that are cheap to rerun together, rather than merging entire independent workstreams.
I would enforce consistent step naming, mandatory comment blocks describing each step's purpose, and standardized condition code thresholds across the shop so on-call staff can reason about any job without reading it line by line first. I would also push for centralized documentation linking each job to its business function and the team responsible for it.
I would review job class and initiator assignments to make sure batch work is properly isolated from time-critical online regions, and look at WLM policy definitions to see if batch service classes need retuning. I would also examine whether any batch jobs could shift into an off-peak window or run with a lower priority without breaking their own SLA, freeing capacity during contention periods.
I would evaluate raising the LIMIT on the GDG base definition against actual retention requirements, since simply raising the limit without archiving old generations just delays the same problem. I would set up an archival or scratch process for generations beyond the retention policy and make sure any downstream job referencing a fixed relative generation number would not break if the limit changes.
I would design the update step to work against a copy or staging dataset first, then commit to the real master only after validation, so a mid-step failure never leaves the master file in a half-updated state. If the update has to happen in place for performance reasons, I would build explicit checkpoint and rollback logic into the program itself, since JCL alone cannot undo a partial file update.
Dynamic generation, where a driver program builds the JCL text at submission time, gives more flexibility for cases with many varying parameters or conditional steps that symbolics cannot express cleanly, but it is harder to audit and version control than a static proc. I would default to symbolic parameters in a cataloged proc whenever the variation is bounded and predictable, and reserve dynamic generation for cases where the step structure itself needs to change based on runtime conditions.
I would look at how much idle wait time currently exists in the cycle because jobs are gated on clock time rather than actual data availability, since event-driven triggering usually recovers that wasted time directly. I would prioritize migrating the streams with the most cross-team dependencies first, since those tend to have the most slack built in defensively around uncertain timing.
I would review the sort utility's SORTWK allocation and make sure enough work datasets with adequate space are available, since undersized sort work areas force extra I/O passes. I would also check whether the sort can run with a larger REGION and whether the data volume justifies switching to a parallel or dynamic allocation sort configuration rather than fixed work dataset counts.
8-10 Years
I would establish a shared procedure library governance model with change review gates, mandatory naming conventions enforced by validation tooling, and clear ownership boundaries so one team's proc change cannot silently break another team's job. I would also require impact analysis before any shared proc modification, since cataloged procedures used across teams create hidden coupling that is easy to break without realizing it.
I would frame the case around the growing cost of scarce mainframe skills, the operational risk of undocumented job dependencies accumulated over decades, and the opportunity cost of batch windows constraining how fast the business can move data downstream. I would pair that with a phased modernization plan that targets the highest-risk, highest-maintenance jobs first rather than proposing a disruptive wholesale rewrite.
I would define a standard scale, such as 0 for success, 4 for warning, 8 and above for hard failure, and require new development to follow it strictly while documenting exceptions in legacy jobs that cannot be safely changed. I would prioritize remediating the legacy jobs whose inconsistent codes actively cause scheduler misfires or missed alerts, rather than trying to retrofit the entire estate at once.
I would assess which jobs depend on true z/OS-specific features like VSAM clustering or JES-level scheduling quirks that an emulator might not replicate faithfully, since those are the highest-risk candidates for subtle behavioral differences after migration. I would push for a pilot with a bounded, well-understood batch stream first, measuring both functional parity and performance under representative volume before committing the wider estate.
I would define a naming hierarchy tied to application ownership and data sensitivity, with retention and disposition rules baked into the standard so DFHSM or equivalent storage management tools can enforce lifecycle policy automatically rather than relying on manual cleanup. I would phase enforcement in through new development first, then require legacy applications to conform at their next major touch point rather than mandating an immediate mass migration.
I would weigh the business's actual latency requirements against the cost and risk of rearchitecting, since not every batch process benefits from real-time conversion and some genuinely belong in a nightly window. I would prioritize candidates where downstream consumers are already asking for fresher data or where the batch window itself is under capacity pressure, and treat the rest as lower priority modernization work.
I would build a template library of pre-approved, parameterized procs covering the common patterns, paired with automated validation tooling that checks naming conventions, space allocations, and return code handling before a job can be promoted. I would keep a lightweight review gate only for jobs that fall outside the standard templates, so routine work does not bottleneck on a central team.
I would quantify the current cost of undiagnosed batch failures in terms of missed SLAs, manual investigation time, and downstream business impact, since that translates technical debt into numbers a non-technical stakeholder can weigh against the tooling's cost. I would also highlight that better observability reduces dependency on a shrinking pool of specialists who currently diagnose issues by memory and tribal knowledge.
I would bring both teams together to separate their actual functional requirements from implementation preferences, and look for whether the shared proc should instead be split into two variants with a common base, since forcing one proc to serve genuinely divergent needs usually produces something fragile for everyone. I would document the decision and its rationale so future teams inheriting the proc understand why it is structured the way it is.
I would work with compliance and the business owners of each data domain to set retention periods based on actual regulatory and operational recovery requirements rather than a single blanket policy, since some data genuinely needs years of generations while other work files need only a few days. I would automate enforcement through storage management tooling so retention decisions do not depend on individual job maintainers remembering to clean up.
I would define a shared minimum standard for what every job must log and how return codes map to alert severity, then work with each team on a migration timeline rather than mandating an overnight rewrite of every job's error handling. I would prioritize business-critical jobs for early conversion so the alerting improvement pays off where it matters most first.
I would look at how much of the JCL across the portfolio follows genuinely repeatable patterns versus how much has legitimate one-off complexity, since generator tooling pays off fastest where there is high repetition and gets awkward when every job is a special case. I would pilot the tooling on the most template-friendly application first to prove the maintenance savings before rolling it out organization wide.
I would restrict production JCL changes to a formal promotion process with peer review and change ticket linkage, while allowing broader access in test regions so developers can iterate quickly without bottlenecking on approvals. I would tie the boundary to actual audit and compliance requirements rather than an arbitrary hierarchy, since overly restrictive test access just pushes developers toward risky workarounds.
I would look at whether the application's current single-LPAR footprint creates a meaningful single point of failure for business-critical processing, weighing that against the operational complexity of coordinating cross-LPAR job dependencies and shared dataset access. I would prioritize splitting only where the business impact of an LPAR-level outage clearly outweighs the added coordination overhead.
10+ Years
I would start them on small, well-bounded changes like adjusting a SPACE parameter or adding a comment block to an existing job, building confidence with the toolchain before exposing them to complex multi-step dependency chains. I would pair that with walking through a few representative jobs together, explaining what each statement does and, beyond just the syntax, why the shop's conventions evolved the way they did.
I would prioritize documenting the dependency graph and business purpose of the cycle first, since that context is the hardest thing to reconstruct once the people who built it are gone, and pair the departing engineers with successors on live troubleshooting sessions rather than relying purely on written documentation. I would also push to record decisions and rationale as they come up, beyond just the final state, since future maintainers need to understand why the cycle is shaped the way it is.
I would translate the technical risk into business terms, framing it around what happens to critical financial or operational processing if the batch cycle fails and few people remain who can diagnose it quickly. I would pair that risk narrative with a concrete, phased investment plan rather than an open-ended ask, since executives respond better to a bounded proposal with milestones than a vague call for more resources.
I would separate rotation duties so engineers spend dedicated time on production support without also being expected to deliver new development in the same sprint, since context switching between firefighting and building tends to degrade both. I would also build a shared runbook culture so on-call knowledge does not live only in the heads of the most senior engineers.
I would invest in documentation, standardized templates, and automated validation tooling that encodes institutional knowledge into repeatable patterns rather than tribal memory, while also building a deliberate pipeline of newer engineers cross-trained on the batch estate. I would pair that with selectively identifying which parts of the estate are genuinely candidates for modernization or retirement, since not every legacy job needs the same level of investment to keep running safely.
I would present a clear picture of what is actually driving the increased complexity, whether it is business growth, technical debt, or both, and show how the proposed cuts map against risk of missed SLAs or extended outages. I would advocate for targeted investment in automation and documentation as a way to do more with a smaller team sustainably, rather than framing it purely as a headcount defense.
I would involve representatives from each team early in drafting the standard rather than issuing it top-down, since standards adopted without buy-in tend to get quietly ignored once the mandate loses momentum. I would also build in a clear rationale for each standard so teams understand the operational problem it solves, beyond just the rule itself.
I would frame the initiative around measurable business outcomes like reduced batch window risk, lower operational cost over time, and improved agility for downstream data consumers, avoiding technical jargon that obscures the actual value being delivered. I would present it in phases with clear go or no-go checkpoints, since a multi-year commitment is easier for leadership to support when they can see incremental proof points along the way.
I would push for investment in dashboards and trend reporting on batch runtime, return code patterns, and near-miss SLA events, since visibility into slow degradation is what allows teams to act before a full outage happens. I would also recognize and reward engineers for catching and fixing latent issues proactively, since a culture that only celebrates heroic incident recovery inadvertently discourages preventative work.
I would identify a designated backup early and structure their onboarding around shadowing real incidents and change reviews rather than passive documentation review alone, since batch troubleshooting skill is built through exposure to real edge cases. I would also formalize the expert's implicit knowledge into runbooks and decision trees as part of the transition, capturing judgment calls that are easy to lose once that person moves on.
I would weigh the frequency of change against the job's criticality, refactoring jobs that get touched often and carry high business risk since messy JCL there compounds maintenance cost over time, while leaving rarely touched stable jobs alone if they are working correctly. I would avoid refactoring purely for aesthetic reasons on jobs that are stable and low risk, since that introduces unnecessary regression risk without a corresponding benefit.
I would be direct about the fact that a large, interdependent batch estate cannot be modernized quickly without unacceptable risk, and propose a realistic phased timeline tied to natural business cycles like quarter-end processing windows where testing and rollback have the least impact. I would keep stakeholders updated with concrete milestones rather than vague progress updates, so trust in the timeline builds as each phase lands successfully.
I would create structured opportunities for promising engineers to lead smaller initiatives, like owning a modernization pilot or a cross-team standards effort, so leadership capability gets demonstrated through real work rather than assumed from years of service. I would pair that with direct, honest feedback on both technical and communication growth areas, since strong batch engineering skill does not automatically translate into leading people or influencing stakeholders.
I would work with both sides to understand the actual risk profile driving each position, then propose a tiered release process where low-risk changes like parameter tuning move faster while higher-risk structural changes get the fuller review cycle. I would frame the compromise around shared incentives, since both teams ultimately want fewer production incidents, just with different assumptions about how to get there.




