pgvector in production: indexing, recall, and the queries that got slow
We built a semantic search and retrieval system for 12 million financial filings and real-time news articles. Our choice of vector database wasn’t a speci…
Read entryWe built a semantic search and retrieval system for 12 million financial filings and real-time news articles. Our choice of vector database wasn’t a speci…
Read entryTwo years ago, my team was responsible for maintaining the production deployment pipelines of a suite of real-time machine learning models. These models p…
Read entryDuring high-volatility market events, such as an unscheduled Federal Reserve rate announcement, my team's algorithmic trading infrastructure processes a m…
Read entryA few months ago, I was tasked with building an automated quantitative research agent. The goal was straightforward: ingest a ticker symbol, pull real-tim…
Read entryWe run a real-time sentiment and market-impact parsing engine that processes news feeds, regulatory filings, and earnings transcripts. Our SLA requires us…
Read entryIn high-throughput, LLM-powered trading intelligence pipelines, latency is the bottleneck that kills execution edge. My team runs an automated pipeline th…
Read entryWe were spending $4,200 a month on closed-source LLM APIs to power a real-time sentiment extraction and limit-order-book feature pipeline. The setup was s…
Read entrySix months ago, our team was running a cluster of 64 NVIDIA H100 GPUs hosting a pipeline of custom Stable Diffusion XL and LLaMA-3-70B models. For the fir…
Read entryI've shipped agent workflows on both LangGraph and Temporal, and I keep seeing them framed as competitors. They're not. They solve different problems, and…
Read entryAt 2:14 AM the pager went off: every API request was timing out with remaining connection slots are reserved for non-replication superuser connections. Th…
Read entry