Introducing My DIT/PhD Research Project – Hybrid Deployments and the Challenge of Observability at Enterprise Scale

As part of my journey through the Doctor of Information Technology (DIT/PhD) program I’m exploring one of the most critical challenges in enterprise cloud computing today: observability in hybrid service mesh environments.

The Core Issue

In today’s hybrid enterprises, platform engineers, SREs, and DevOps teams are expected to ensure uptime, trace issues fast, and report on system health – across cloud and on-prem.

But when things break?
They face the same frustrating reality:

  • Disconnected logs and telemetry

    • (hybrid deployments)

  • Manual firefighting during incidents

    • (inter-connecting hybrid deployments)

  • Gaps in visibility across environments

    • (hybrid deployments inter-connectivity)

This hits hard during:

  • Incident reviews (longer resolution times)

  • Executive reporting (limited visibility)

  • Tool and budget decisions (unclear ROI)

What My Research Explores

My capstone project aims to develop a practical framework that helps large enterprises:

  • Unify observability practices across hybrid deployments

  • Reduce incident resolution times

  • Help IT leaders make smarter data-driven decisions during executive IT reporting, annual planning meetings, and next-year budget reviews

This is a qualitative study informed by:

  • Practitioner interviews

  • Real-world hybrid cloud pain points

  • And guided by Systems Thinking (Checkland, 1999)

Why Now?

Hybrid cloud is no longer optional.
Without structured observability, enterprises:

  • Burn time during incidents

  • Overspend on tools

  • Underdeliver on reliability

What’s Next

Over the coming weeks, I’ll:

  • Interview practitioners

  • Analyze patterns

  • Share findings as I build a framework that works in the real world

If you lead platform teams, care about observability, or are scaling hybrid infrastructure – follow along. Let’s fix the visibility gap where it matters most.

References

Chelliah, P. R., & Surianarayanan, C. (2021). Multi-cloud adoption challenges for the cloud-native era: Best practices and solution approaches. International Journal of Cloud Applications and Computing, 11(2), 67–96. https://doi.org/10.4018/IJCAC.2021040105

Duan, Y., Bao, H., Bai, G., Wei, Y., Xue, K., You, Z., Zhang, Y., Liu, B., Chen, J., Wang, S., & Ou, Z. (2024). Learning to diagnose: Meta-learning for efficient adaptation in few-shot AIOps scenarios. Electronics, 13(11), 2102. https://doi.org/10.3390/electronics13112102

Poghosyan, A., Harutyunyan, A., Davtyan, E., Petrosyan, K., & Baloian, N. (2024). The diagnosis-effective sampling of application traces. Applied Sciences, 14(13), 5779. https://doi.org/10.3390/app14135779

Valli, L. N., Sujatha, N., Mech, M., & Lokesh, V. S. (2024). Accelerate IT and IoT with AIOps and observability. E3S Web of Conferences, 491, 4021. https://doi.org/10.1051/e3sconf/202449104021

Zhang, D., & Zheng, M. (2021). Benchmarking for observability: The case of diagnosing storage failures. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 1(1), 100006. https://doi.org/10.1016/j.tbench.2021.100006

Checkland, P. (1999). Systems thinking, systems practice. John Wiley & Sons. https://www.wiley.com/en-us/Systems+Thinking%2C+Systems+Practice%3A+Includes+a+30-Year+Retrospective-p-9780471986065