Skip to content
View ai-hpc's full-sized avatar
🦙
Research Inference Optimization & Confidential Computing
🦙
Research Inference Optimization & Confidential Computing

Highlights

  • Pro

Organizations

@openclaw @FastCrest @cathedralai @gittensor-model-hub

Block or report ai-hpc

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ai-hpc/README.md

GPU Inference, AI Hardware & Agent Systems Engineer

CUDA/C++ · Blackwell LLM Inference · Agent Systems · Confidential Computing


I build high-performance AI systems across the complete hardware–software stack — from GPU kernels and inference runtimes to optimized models, autonomous agents, and verifiable compute infrastructure.

My foundation is in embedded Linux, firmware, real-time systems, hardware design, and silicon bring-up. My current work focuses on Blackwell-native LLM inference, model optimization, agent systems, and confidential computing.


Current Work

  • SparkInfer — native C++/CUDA inference runtime optimized for NVIDIA Blackwell GPUs

  • LLM Performance Engineering — NVFP4/FP8 quantization, speculative decoding, long-context inference, KV-cache optimization, tool calling, and RTX 5090 deployment

  • Agent Systems — personal AI agents connecting models, tools, skills, memory, communication channels, and device-local actions

  • Confidential & Verifiable Computing — trusted execution, GPU attestation, signed compute evidence, fail-closed validation, sealed evaluation, and verifiable sandbox infrastructure

  • Physical & Edge AI — Jetson, voice and perception pipelines, embedded Linux, IoT, local automation, and hardware-integrated intelligence


Engineering Focus

  • CUDA/C++ kernel and inference-runtime optimization
  • NVIDIA Blackwell architecture and quantized inference
  • LLM serving, scheduling, memory management, and latency optimization
  • Long-context inference and speculative decoding
  • Agent runtimes, tool execution, memory, and multi-step workflows
  • Confidential computing, TEE/GPU attestation, and proof-backed execution
  • Embedded Linux, firmware, and hardware–software co-design

Building

A complete AI systems stack connecting:

hardware → kernels → runtime → optimized models → agents → verified execution → physical systems

My long-term goal is to design a custom inference accelerator optimized for agent workloads.


Fast inference. Autonomous agents. Verifiable compute.

Pinned Loading

  1. ai-hardware-engineer-roadmap ai-hardware-engineer-roadmap Public

    Master AI inference, AI agent harness systems, and hardware engineering — then design a physical AI chip. That is the goal.

    HTML 258 37

  2. awesome-ai-hardware awesome-ai-hardware Public

    AI accelerators, edge inference devices, compilers, runtimes, benchmarks, and research for building and evaluating machine-learning systems.

    18 5

  3. gittensor-ai-lab/sparkinfer gittensor-ai-lab/sparkinfer Public

    Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.

    C++ 15 68

  4. GeniePod/genie-ai-runtime GeniePod/genie-ai-runtime Public

    Jetson Orin-tuned LLM inference runtime for gemma 4, qwen 3.5 — memory-first, power-aware, zero-allocation. C++17 + CUDA.

    Cuda 1 2

  5. GeniePod/genie-claw GeniePod/genie-claw Public

    🦞 Low-latency, limited-context AI harness for private on-device homes.

    Rust 48 70

  6. neptuneprivacy/triton-vm-prover neptuneprivacy/triton-vm-prover Public

    High-performance C++/CUDA GPU-accelerated STARK prover for Triton VM

    Rust 4 1