Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
aiengineering
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
WebGPU LLM Inference: Running 7B Models Natively in the Browser
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
WebGPU LLM Inference: Running 7B Models Natively in the Browser
#
aiengineering
#
webgpu
Comments
Add Comment
4 min read
Serving Mixture of Experts (MoE): Memory-Efficient Inference Routing
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
Serving Mixture of Experts (MoE): Memory-Efficient Inference Routing
#
aiengineering
#
architecture
Comments
Add Comment
3 min read
S-LoRA: Multiplexing Thousands of Fine-Tuned Adapters on a Single GPU
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
S-LoRA: Multiplexing Thousands of Fine-Tuned Adapters on a Single GPU
#
aiengineering
#
performance
Comments
Add Comment
3 min read
DSPy: Replacing Prompt Engineering with Declarative Optimization Compilers
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
DSPy: Replacing Prompt Engineering with Declarative Optimization Compilers
#
aiengineering
#
prompting
Comments
Add Comment
3 min read
Structured Output Generation: Enforcing JSON & Regex at the Logits Level
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
Structured Output Generation: Enforcing JSON & Regex at the Logits Level
#
aiengineering
#
logits
Comments
Add Comment
3 min read
ColBERT Late Interaction: Advancing RAG Beyond Dense Embeddings
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
ColBERT Late Interaction: Advancing RAG Beyond Dense Embeddings
#
aiengineering
#
rag
Comments
Add Comment
4 min read
KV Cache INT4 Quantization for 1M+ Token Context Windows
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
KV Cache INT4 Quantization for 1M+ Token Context Windows
#
aiengineering
#
quantization
Comments
Add Comment
3 min read
OmniRouter Architecture: Resilient LLM Gateway Routing & Fallback Pipelines
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
OmniRouter Architecture: Resilient LLM Gateway Routing & Fallback Pipelines
#
aiengineering
#
architecture
Comments
Add Comment
5 min read
Multi-Agent Swarm Orchestration: Hierarchical Agentic Workflows
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
Multi-Agent Swarm Orchestration: Hierarchical Agentic Workflows
#
aiengineering
#
agents
Comments
Add Comment
3 min read
vLLM PagedAttention: Memory Optimization & High-Throughput LLM Inference Tuning
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 14
vLLM PagedAttention: Memory Optimization & High-Throughput LLM Inference Tuning
#
aiengineering
#
performance
Comments
Add Comment
4 min read
An Agent Is a Loop: a Working Mental Model for Agentic Systems
Xinyang Wu
Xinyang Wu
Xinyang Wu
Follow
Aug 3
An Agent Is a Loop: a Working Mental Model for Agentic Systems
#
llm
#
agents
#
architecture
#
aiengineering
Comments
1
 comment
6 min read
Logistic Regression Doesn't Make Decisions—Your Business Does
Nishant Banginwar
Nishant Banginwar
Nishant Banginwar
Follow
Jul 30
Logistic Regression Doesn't Make Decisions—Your Business Does
#
machinelearning
#
aiengineering
#
mlops
#
sre
Comments
Add Comment
3 min read
Best Practices for AI-DLC: Amazon
ke yi
ke yi
ke yi
Follow
Jul 24
Best Practices for AI-DLC: Amazon
#
aiengineering
#
developerproductivity
Comments
Add Comment
17 min read
Building a Pi Agent from Scratch (7-Day Retrospective)
Scc_hy
Scc_hy
Scc_hy
Follow
Aug 5
Building a Pi Agent from Scratch (7-Day Retrospective)
#
llm
#
agents
#
aiengineering
#
contextmanagement
Comments
3
 comments
29 min read
Your banned-word list expired in March 2024
Alex Chernysh
Alex Chernysh
Alex Chernysh
Follow
Aug 12
Your banned-word list expired in March 2024
#
writing
#
editing
#
agentskills
#
aiengineering
Comments
1
 comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account