Skip to content
View asher0913's full-sized avatar
  • University of Illinois
  • Champaign, IL

Highlights

  • Pro

Block or report asher0913

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
asher0913/README.md

Yixuan Zhang

I build machine-learning systems that hold up outside the notebook: private split inference, edge vision and LLM serving. I'm an MS CS student at the University of Illinois Urbana-Champaign ('28), after a First-Class B.Sc. in Computer Science at the University of Nottingham.

Website · Résumé · LinkedIn

Start here

If you are hiring for Look at What to check
Software engineering multi-material-slicer, code-task-forge a C++17/Qt/OpenGL desktop app with a headless import-to-G-code self-test in CI; a shell-free, time-limited patch runner behind a CLI, FastAPI and Docker
Machine learning engineering DualPathCEM, edge-quantization-lab, recsys-ranking-lab an audited research snapshot; quantization with ./scripts/demo.sh rerunning every number and claim; a two-stage recommender on MovieLens-100K
LLM systems and agents inference-lab, secure-rag, skill-router batching, coalescing and semantic caching for LLM serving; permission checks in RAG; tool routing measured on held-out phrasing

Each README says which results come from real data, synthetic data or simulation, and which of them CI reruns.

Research

DualPathCEM Split inference for private face recognition. Two client paths: a noise-protected spatial tensor carries the privacy burden and a compact semantic token carries utility. 81.96% top-1 on FaceScrub, +1.63 points over Noise_ARL+CEM, with 75.8% higher reconstruction error for a decoder attacker. Manuscript in preparation.
SlotCEM B.Sc. dissertation (First Class). Slot-attention conditional-entropy regularisation against model inversion in split learning: +29.8% attack MSE for a 0.84-point accuracy cost.

Selected work

multi-material-slicer Internship project. Qt/C++17 and OpenGL desktop slicer for multi-material resin printing: STEP/STL assemblies to per-material masks and G-code, with a headless end-to-end self-test in CI.
skill-router Progressive tool disclosure for large skill catalogues: −85.5% context tokens per query. Measured on held-out paraphrases, where lexical routing drops to 0.21 top-1 and an embedding stage brings it back to 0.56.
inference-lab LLM serving front end for vLLM, SGLang and Ollama. Request coalescing cuts backend generations 16 → 3. A labelled study shows semantic caches serving wrong answers, and a simulation shows continuous batching sustaining 10× the load of static batching.
secure-rag Permission-aware multi-tenant RAG. After ACL changes, a stale-index pre-filter leaks on 54% of queries; a live re-check against the directory of record brings that to 0%.
code-task-forge SWE-bench-style harness for judging coding-agent patches. Exit codes accept 25 of 78 wrong patches; restored tests plus held-out tests accept none.
plant-leaf-recognition ResNet-101 and ViT-B/16 feature fusion: 98.46% top-1 on 100 leaf species, reported from the original run over 20 unstratified splits.
oxford-pet-classification 37 breeds: fine-tuned ResNet-18 at 89.8% against a from-scratch SE-ResNet at 49.7%, with a one-change-per-row ablation and Grad-CAM.
IAMABOT MechMania 32 bot, built from the game engine's Rust source; climbed from 7th to 3rd on the live leaderboard.

Labs

Laptop-scale reference implementations of production problems, each with seeded, reproducible results checked in CI. Synthetic benchmarks are labelled as such in each README.

Serving and systems

  • tiered-kv-cache-lab: KV prefix caching across HBM, DRAM and NVMe, with capacity planning up to 14.5 req/s within the TTFT SLO.
  • ai-gateway-control-plane: budgets, circuit breakers and latency-aware routing; p99 latency 14.3 s → 5.8 s under fault injection.
  • edge-quantization-lab: INT8/INT4 post-training quantization. The mixed-precision search reaches 6.7× compression and matches the exhaustive optimum on all 5 seeds. An integer-only kernel agrees with simulation on every prediction.
  • vision-pipeline-orchestrator: CPU/GPU pipeline simulator with batching, backpressure, a dead-letter queue and autoscaling.

Agents and evaluation

  • agentops: incident investigation under flaky telemetry: wrong on 0.4% of incidents against 33.5% for retry-and-answer, with replayable traces.
  • agent-trace-lab: trace model, process anti-pattern detectors and root-cause attribution, stress-tested on 1,000 generated traces.
  • reflective-agent-lab: reflect-and-repair against sandbox verification; irreversible side effects fall from 12.0 to 0.2 per 100 tasks.
  • graph-sop-agent: service graph plus runbooks, against document-only RAG, with review-gated knowledge ingestion.
  • risk-agent-workbench: span-cited evidence extraction that abstains when evidence is missing.
  • agent-guard: policy gateway between an agent and its tools.

Retrieval, data and safety

Machine learning and algorithms

Apps


Python · C++ · Java · Swift · SQL · PyTorch · scikit-learn · FastAPI · Qt/OpenGL · Docker · GitHub Actions

Pinned Loading

  1. DualPathCEM DualPathCEM Public

    Dual-path split inference for private face recognition: 81.96% top-1 on FaceScrub (+1.63 pts over Noise_ARL+CEM) with 75.8% higher decoder-attack reconstruction error. Frozen snapshot of a manuscri…

    Python

  2. multi-material-slicer multi-material-slicer Public

    Qt/C++17 desktop slicer for multi-material resin printing: STL/STEP assemblies, OpenGL preview, per-material PNG masks, DP tank-change planning and G-code, with a headless import-to-G-code self-tes…

    C++

  3. inference-lab inference-lab Public

    LLM serving front end for vLLM/SGLang/Ollama: tenant/model/limit-keyed semantic cache, single-flight coalescing (16 → 3 backend generations), per-request limits in dynamic batches; a labelled study…

    Python

  4. code-task-forge code-task-forge Public

    SWE-bench-style harness for judging coding-agent patches: path policy, shell-free runs in throwaway workspaces, restored tests, fail-to-pass/pass-to-pass, held-out tests and a failure taxonomy. Exi…

    Python

  5. secure-rag secure-rag Public

    Permission-aware multi-tenant RAG: signed-token identity, directory groups, fail-closed ACLs, pre-filtered ranking and a live re-check that withholds revoked or stale text. A stale-index pre-filter…

    Python

  6. edge-quantization-lab edge-quantization-lab Public

    Post-training quantization from first principles: INT8/INT4, per-channel scales, calibration observers, byte-aware mixed-precision search (6.7×, matches the exhaustive optimum on 5/5 seeds) and an …

    Python