
Supporting Complex Customer Environments
Diagnosing production issues across cloud infrastructure, customer networks, and on-premise systems
Background
Although ProShop is primarily delivered as a cloud-hosted SaaS platform, some customers choose to deploy portions of the system within their own infrastructure. Depending on operational requirements, this may include customer-hosted virtual machines, local file storage, private network resources, or integrations that require secure communication between cloud-hosted services and systems operating inside the customer’s network.
To support these deployments, ProShop provides a pre-configured virtual machine that gives customers a consistent starting point. However, every customer environment evolves differently over time. Infrastructure, networking, security policies, operating system configuration, virtualization platforms, and third-party software all contribute to unique production environments.
As a result, no two customer deployments are exactly alike.
The Challenge
Supporting customer-managed infrastructure presents a very different set of challenges than operating a traditional SaaS platform.
Issues that appear to originate within the application are often influenced by factors outside of it. Network connectivity, DNS, firewalls, virtualization, operating system configuration, permissions, storage, or third-party software can all affect application behaviour.
Many problems cannot be reproduced internally because they only occur within the customer’s production environment. Diagnosing these issues requires understanding how the entire system fits together rather than focusing on a single application or service.
The challenge is not simply fixing software defects—it’s determining where the problem actually exists.
My Approach
When investigating production issues, I start by building an understanding of the customer’s environment before drawing conclusions about the application itself.
Rather than assuming the application is at fault, I work systematically across application logs, cloud infrastructure, networking, operating systems, and customer-managed resources to isolate the source of the problem. In many cases, this means collaborating with implementation teams, customer IT staff, and support engineers to eliminate potential causes until the underlying issue becomes clear.
Over time, this work has strengthened my ability to troubleshoot complex systems that span cloud services, on-premise infrastructure, and third-party integrations.
My Role
A significant part of my day-to-day role involves supporting customers whose environments extend beyond the boundaries of our SaaS platform.
I regularly investigate production issues involving cloud infrastructure, Windows and Linux systems, networking, virtual machines, APIs, customer-hosted resources, and third-party integrations. Because every environment is different, each investigation requires adapting to unique configurations rather than relying on predefined solutions.
What I enjoy most about this work is that every issue is different. The process of gathering evidence, forming hypotheses, and tracing problems across multiple layers of the technology stack keeps the work both challenging and rewarding.
The Outcome
Successfully supporting hybrid customer environments requires more than technical knowledge—it requires the ability to understand how independent systems interact in real-world production environments.
By approaching problems methodically and considering the entire ecosystem rather than a single application, I have helped resolve complex customer issues, improve production reliability, and support successful deployments across a wide variety of customer environments.
Lessons Learned
This aspect of my role has reinforced that the most difficult production issues rarely have a single obvious cause. They often emerge from the interaction between applications, infrastructure, networks, operating systems, and customer-specific configuration.
It has also taught me that effective troubleshooting is as much about understanding systems as it is about understanding software. The ability to identify the true root cause—especially in environments that cannot be reproduced internally—is one of the most valuable engineering skills I’ve developed throughout my career.
Technologies
- AWS Elastic Container Service (ECS)
- Linux
- Golang, gRPC, Protocol Buffers
- QuickBooks, SAGE50
Key Takeaway: I learned that solving production issues is usually less about knowing the answer and more about understanding how all the pieces fit together.
