Module 13: CICS Recovery, Restart and Production Support
Common Production Issues
Production CICS regions fail in familiar ways. This page lists the issues every CICS support person meets sooner or later, how to recognize them, and the first thing to do about each.
Storage and capacity problems
- Short on storage (SOS) - the region runs out of dynamic storage; new tasks fail and CICS sheds load. Find the storage hog with dumps or statistics.
- MAXTASKS reached - no new tasks can start and users see delays. Check for hung or looping tasks holding task slots.
- DSA limit - the dynamic storage area is sized too small for the workload. Review DSA usage in the statistics and enlarge it.
- Storage violation - a program overwrote CICS control blocks. The violating program must be found from the dump and fixed.
Task and performance problems
- Runaway tasks (AICA) - infinite loops burn CPU and block a task slot. Purge the task and fix the loop.
- Deadlocks (ATCH/AKCS) - tasks wait on each other's locks until CICS kills one. Shorten units of work and fix update ordering.
- Slow response times - check for file I/O waits, DB2 delays, or too few MAXTASKS before blaming the network.
- Dump datasets full - transaction dumps stop being taken and diagnosis becomes impossible. Offload or enlarge the dump datasets.
Definition and resource problems
- Programs, files or transactions coming up disabled after a change - a bad definition was installed. Back it out and reinstall the previous version.
- New program version not picked up - the old copy is still loaded. Use CEMT SET PROGRAM(...) NEWCOPY (or PHASEIN) to load the new one.
- Journals filling up - recovery logging stalls the region. Monitor journal usage and offload on schedule.
