Machine Learning Engineer at Amazon · Building assevra.ai · IEEE-published researcher

Agentic AI & Distributed Systems Engineer

I build reliable, large-scale agentic AI and distributed systems — and research how to make autonomous software trustworthy in production. Eleven-plus years spanning enterprise B2B commerce, cloud platforms, and low-latency trading.

Portrait of Veera Ravindra Divi

What I work on

Professional focus

  • Agentic & Autonomous AI

    LLM orchestration, multi-agent architectures, and developer copilots that plan, act, and self-improve within governed boundaries.

  • Trustworthy & Secure GenAI

    Reliability scoring, guardrails, evaluation harnesses, and hallucination-risk validation for autonomous systems in production.

  • Distributed & Event-Driven Systems

    High-throughput, low-latency, fault-tolerant platforms built on resilient messaging and cloud-native infrastructure.

  • B2B Commerce & eProcurement

    Enterprise ordering, ERP and marketplace integration, multi-entity procurement, and resilient supply-chain execution.

  • AI-Assisted Software Engineering

    Autonomous CI/CD, release-safety automation, and AI-augmented code review that raise developer velocity without lowering the bar.

  • Cloud-Native Platforms

    API design, service reliability, and operational and cost optimization across large-scale AWS environments.

Selected research

Publications

  • publishedIEEE · 2026 5th International Conference on Electronics Representation and Algorithm (ICERA), Yogyakarta, Indonesia

    GuardGRPC: Priority-Aware Resource Governance for OTLP/gRPC Telemetry Export

    A priority-aware governance layer for OTLP/gRPC export that adapts telemetry controls under pressure while preserving incident-critical signals.

  • publishedIEEE · 2026 5th International Conference on Electronics Representation and Algorithm (ICERA), Yogyakarta, Indonesia

    Containing the Cascade: A Benchmark and Reference Mediator for Failure Propagation in Tool-Using LLM Multi-Agent Systems

    A taxonomy of cascading failures, the Cascade-Bench fault-injection testbed, and Sentinel, a reference mediator for containing failures in tool-using multi-agent systems.

  • publishedIEEE · 2026 Seventeenth International Conference on Ubiquitous and Future Networks (ICUFN), Milan, Italy

    SecReviewAgent: Context-Aware Security Review of Infrastructure-as-Code Using Persistent Architecture Memory

    An LLM-powered Infrastructure-as-Code security-review system that uses persistent architecture memory to evaluate changes in repository context.

  • publishedIEEE · 2026 International Conference on Artificial Intelligence, Systems, and Emerging Technologies (ICAISET), Cairo, Egypt

    CostAgent: Self-Improving Autonomous LLM-Based Orchestration for Cost-Optimal Cloud Data Processing at Scale

    A bounded-risk autonomous orchestration framework for cost-aware cloud data processing on preemptible infrastructure with deterministic safety enforcement.

All research & publications →

Open source

Selected project

Assevra AI

An MIT-licensed, open-source reliability scorecard for LLM agents. Assevra AI scores agent reliability against fixed thresholds so teams can measure — not guess — whether an agent is safe to ship.

  • Python
  • LLM evaluation
  • CLI
  • CI/CD

All projects →

Let's talk

Open to research collaboration, speaking, technical advising, and conversations on agentic AI reliability.

Get in touch