Skip to content
View SayedAbbas's full-sized avatar
🤘
🤘

Block or report SayedAbbas

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SayedAbbas/README.md

Shabi Abbas Sayed

Senior Applied AI Solutions Architect @ AWS · 2× AWS re:Invent Speaker

I help enterprises turn ambitious AI ideas into production systems. My focus is production-grade Enterprise AI Agents: agents with measurable quality, grounded responses, controlled tool access, and clear operational boundaries.

At a glance

  • Business-first Applied AI — start with the customer problem and measurable outcome, then decide where AI adds value.
  • Hands-on architecture — agents, tool calling, MCP, RAG, realtime/Voice AI, evals, guardrails and enterprise integrations.
  • Production discipline — deterministic controls around high-impact actions, explicit human approval, observability and regression evaluation.
  • Public proof — runnable projects, CI-backed evaluation harnesses, architecture patterns and production-focused technical writing.

Start here

If you're looking for... Start with
Mistral + MCP change-risk architecture Mistral Enterprise Change Risk Agent — evidence-grounded production-change investigation with MCP tools, deterministic release gates and human approval.
Jev typed decisions & model routing Jev Enterprise Decision Eval Lab — compare Jev with OpenAI, Claude and Gemini across decision quality, calibration, latency, cost and confidence-gated escalation.
Production agent evaluation Enterprise Agent Eval Lab — compare frontier models on business outcomes, grounding, tool use, safety, escalation and latency.
Grok / xAI applied to an enterprise workflow Managed Services SLA Recovery Agent — Grok-powered evidence synthesis with deterministic SLA math and approval boundaries.
AI tied to a commercial outcome RevenueGuard — evidence-backed revenue leakage investigation with deterministic reconciliation and bounded Claude tool use.
Voice AI projects and architectures Explore the Voice AI section.
Production safety & governance Production Applied AI — evals, Responsible AI, agentic governance and production patterns.

How I think about production AI

Start with the business outcome. Give the model room to reason where the path is variable. Keep authorization, policy and high-impact controls deterministic. Then evaluate the complete trajectory—not just the final answer.

What I focus on

  • Enterprise AI Agents & Agent Evals — evaluate task outcomes, tool correctness, grounding, safety, and escalation.
  • Model Context Protocol (MCP) — connect agents to enterprise tools and context with deliberate access boundaries.
  • Frontier Models & Applied AI — design production patterns across Jev, Grok, Claude, OpenAI and other frontier models based on workload fit.
  • Voice / Realtime AI — build conversational experiences that combine realtime interaction, reasoning, enterprise tools and measurable business outcomes.
  • Retrieval-Augmented Generation (RAG) & Guardrails — ground responses in evidence and define when an agent should defer.
  • Production AI — bring reliability, observability, governance, evaluation and operational discipline to applied AI systems.

Featured projects

Project What to explore
Jev Enterprise Decision Eval Lab Compare Jev with OpenAI, Claude and Gemini across enterprise routing, change risk, fraud, claims and support—with calibration, latency, cost, coverage and confidence-based escalation metrics.
RevenueGuard A Claude-powered revenue leakage investigation prototype with deterministic contract, usage, and invoice reconciliation, evidence-linked reviews, synthetic enterprise scenarios, and an offline demo.
Claims Intelligence OpenAI on Amazon Bedrock for evidence-grounded claims review, with explicit human decision boundaries.
SAR Narrative Agent A Claude-based, evidence-grounded suspicious activity report workflow with citations and human review.
Enterprise Agent Eval Lab A vendor-neutral evaluation harness for enterprise agents, with synthetic cases, provider comparisons, and explicit failure signals across task success, grounding, tool use, safety, escalation, and latency.
Managed Services SLA Recovery Agent A Grok-powered agent project focused on predicting enterprise managed-services SLA breaches and recovery workflows.
Mistral Enterprise Change Risk Agent A Mistral + MCP agent for production-change risk: governed evidence tools, traceable evidence IDs, deterministic escalation controls, human approval, evals and CI-backed MCP integration testing.

Try the Eval Lab: v0.2.0 release · No-API-key demo · Sample mock report. The mock uses expected answers to demonstrate the pipeline; its scores are not live-model benchmark results.

Voice AI

I design voice agents around the full customer journey: low latency and natural turn-taking, reliable enterprise tools, clear authorization, human handoff, and evaluation of the action taken. My published work includes these projects and architecture guides.

Projects and AWS Samples

Project What to explore
Grok Voice Enterprise Action Agent Flagship voice project: a runnable, synthetic telecom journey from conversation to diagnostics and controlled action, with Grok Voice, tool calls, deterministic authorization, trajectory evals, and CI.
Claims Intelligence — provider voice assistant An interruptible OpenAI Realtime API voice surface over the same synthetic claim as the specialist workspace. The voice agent can explain status and request documents; claim decisions stay with a human.
Multilingual Voice AI Helpdesk with Amazon Connect A multilingual support experience using Amazon Connect.
Amazon Connect Realtime Analytics and Language Translation Realtime contact center insights and language translation.
Voice AI Assistant with Amazon Nova Sonic An Amazon Nova Sonic voice assistant sample.
Multi-region Resilient Contact Center for Amazon Connect Contact center resilience architecture that supports reliable voice operations.

Production voice architecture writing — in Production Applied AI:

Production Applied AI — Articles

I write practical notes on architectures, patterns, and lessons for moving enterprise AI from prototype to production.

Engineering perspective

An agent's answer is only part of the story. I care about whether it uses the right evidence, selects the right tool, respects permissions, and knows when to escalate. Evaluation should make those behaviors visible before deployment and throughout operation.

If you work on enterprise agents, evaluation, MCP, or realtime AI, explore the projects and join the discussion through repository issues.


Personal projects and views are my own. AWS Samples repositories are maintained under their respective project guidelines and licenses.

Pinned Loading

  1. aws-samples/sample-voice-ai-multilingual-helpdesk-amazon-connect aws-samples/sample-voice-ai-multilingual-helpdesk-amazon-connect Public

    Sample multi-language voice helpdesk on Amazon Connect — routes IT, Payroll, and HR calls in English, Brazilian Portuguese, Mandarin (Putonghua), and Japanese using Amazon Lex and an AWS Lambda + A…

    Python 2 2

  2. aws-samples/sample-amazon-connect-realtime-analytics-and-language-translation aws-samples/sample-amazon-connect-realtime-analytics-and-language-translation Public

    Python 1

  3. aws-samples/sample-voice-AI-assistant-with-amazon-nova-sonic aws-samples/sample-voice-AI-assistant-with-amazon-nova-sonic Public

    Python 1 1

  4. revenueguard revenueguard Public

    Claude-powered enterprise revenue leakage investigation: deterministic billing reconciliation, evidence-linked reviews, and synthetic business scenarios.

    Python 3