Naveen Vivek

Singapore · Agentic systems & AI infrastructure

Naveen KumarVivekanandan

Senior AI Engineer, Dell Technologies

I build agents that take operators from telemetry to diagnosis to controlled, reversible action. Sandboxed, grounded in source and verified before they act.

PlatformsInferenceSandboxesAgents

The journey

Where I am now, and how I got here

  1. Now · Mar 2026 to present

    Dell Technologies

    Senior AI Engineer

    • Digital Twin AI Ops: multi-agent remediation with safe, gated execution
    • MCP Triage Loop, designed and presented to Dell engineers
    • Inference tuning and AI-factory automation on the Dell Automation Platform
  2. Oct 2022 to Mar 2026

    Crédit Agricole

    Big Data Engineer & AI Agent Integrator

    • Petabyte-scale payments and a Google Cloud proof of concept
    • GenAI adoption for 200+ developers
    • Hackathon-winning Code Healer agent
  3. 2013 to 2022

    JP Morgan · Oracle · TCS

    Associate VP, Technical Analyst, Systems Engineer

    • Java and Spring Boot services, Kafka, Apigee and CI/CD
    • 10x load performance at JP Morgan with bulk cache loading
    • Core banking systems for HDFC Bank and State Bank of India

Selected work

Agents that act, safely

What I'm building now, and the work that led to it. Newest first.

Dell Technologies · Current build

Digital Twin AI Ops

A live digital twin of the server estate. Telemetry becomes a graph of servers and their connections, so when something fails the blast radius is already known. Agents reason over it, recommend a fix with evidence, and close the loop.

  1. SenseTelemetry
  2. NormaliseCanonical model
  3. ModelDigital twin
  4. ReasonMulti-agent
  5. ProposeFix + evidence
  6. AuthoriseHuman approval
  7. ExecuteTyped action
  8. VerifyCritic + rollback

Learn. Verified outcomes feed semantic memory, so the next incident starts from what worked last time.

  • Apache Flink
  • Neo4j
  • ClickHouse
  • MongoDB
  • Qdrant
  • OPA

Dell Technologies · June 2026 · Presented to Dell engineers

MCP Triage Loop: from test failure to automated resolution

An agent loop that takes a failing test on a live cluster to a verified fix, a tracked issue or a root-cause fix, with a human reviewing before anything merges.

Two ways in

A failed GitHub Actions build

A developer's manual report

Three steps

Replicate with Playwright MCP

Analyse code and recent commits

Classify the root cause

Three outcomes

Code, config or test fix, then a PR for review

Infrastructure issue, tracked and reported

Flaky test, fixed at the root, never retried away

  • Playwright MCP
  • GitHub MCP
  • Harness
  • GitHub Actions

Earlier work

All projects, including home-lab experiments

How I work

Fast with AI, careful with what it produces

Writing

Latest articles

The first article is on its way

It's about sandboxing the code AI agents write. Follow along on LinkedIn or subscribe to the RSS feed.

On my own time

Learning how the models actually work

Exploring now

  • Deep agents: planner, sub-agents and persistent memory
  • System One models like Jev
  • Google ADK and GKE Agent Sandbox

Toolkit and credentials

What I build with

AI and agents

  • LangGraph
  • LangChain
  • MCP
  • RAG
  • Qdrant
  • Bedrock

Data and streaming

  • Kafka
  • Flink
  • Spark
  • ClickHouse
  • Neo4j
  • JanusGraph
  • MongoDB
  • OpenSearch

Cloud and platform

  • Kubernetes
  • Docker
  • Harness
  • GitHub Actions
  • Bigtable
  • Dataflow
  • Dataproc
  • Playwright

Languages

  • Python (async, FastAPI, Pydantic)
  • Java / Spring Boot
  • REST, gRPC, WebSockets

Building agents that touch real systems?

I'm always happy to compare notes on agent safety, sandboxing and inference.