<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://pvelua.github.io/feed/data.xml" rel="self" type="application/atom+xml" /><link href="https://pvelua.github.io/" rel="alternate" type="text/html" /><updated>2026-10-05T10:45:02-07:00</updated><id>https://pvelua.github.io/feed/data.xml</id><title type="html">Igor’s KB Site | Data</title><subtitle>A personal website for information sharing</subtitle><entry><title type="html">Data and Orchestration Weekly — 4 October 2026</title><link href="https://pvelua.github.io/news/data/2026/10/04/data-weekly/" rel="alternate" type="text/html" title="Data and Orchestration Weekly — 4 October 2026" /><published>2026-10-04T00:00:00-07:00</published><updated>2026-10-04T00:00:00-07:00</updated><id>https://pvelua.github.io/news/data/2026/10/04/data-weekly</id><content type="html" xml:base="https://pvelua.github.io/news/data/2026/10/04/data-weekly/"><![CDATA[<h3 id="langchains-model-router-cuts-coding-agent-cost-by-about-two-thirds-with-no-measurable-quality-loss"><a href="https://www.langchain.com/blog/how-to-build-a-model-router-in-the-harness">LangChain’s model router cuts coding-agent cost by about two thirds with no measurable quality loss</a></h3>

<p><strong>LangChain</strong> · 1 Oct 2026 · <em>Model routing</em></p>

<p>LangChain built a router into the harness of its open-source Open SWE coding agent: a small decision model classifies each new thread once, at the start, and sends it to a fast, balanced or high-performance model tier, which then handles the whole thread. In an A/B test over 973 threads, median cost per thread fell 64% against always using the strongest model, with roughly a third of threads going to the cheapest tier and one in ten to the top tier. Merged-PR rates were statistically indistinguishable between the routed and control groups. It is a concrete, measured case for treating model choice as a middleware concern inside an agent rather than a fixed setting.</p>

<h3 id="qdrant-previews-embedding-models-that-let-you-change-the-query-encoder-without-re-embedding-documents"><a href="https://qdrant.tech/blog/constella-research-preview/">Qdrant previews embedding models that let you change the query encoder without re-embedding documents</a></h3>

<p><strong>Qdrant</strong> · 29 Sep 2026 · <em>Retrieval</em></p>

<p>The Constella research preview encodes documents once with a 400M-parameter model, then offers three query-side options searching those same stored vectors: a token-lookup variant with almost no compute, a 34.5M-parameter transformer, and the full model. Across 15 BEIR datasets the mid-sized option reaches about 91% of the full model’s average score at roughly twelve times the speed, while the lookup variant runs around 480 times faster on a laptop CPU. Models are on Hugging Face with FastEmbed support on a preview branch only, pending internal review before a full release. If it holds up, one index could serve both cheap edge queries and high-quality server queries.</p>

<h3 id="temporal-describes-an-internal-tool-that-orchestrates-security-fixes-across-dozens-of-repositories"><a href="https://temporal.io/blog/camper-running-security-campaigns-on-temporal">Temporal describes an internal tool that orchestrates security fixes across dozens of repositories</a></h3>

<p><strong>Temporal</strong> · 1 Oct 2026 · <em>Durable execution</em></p>

<p>Temporal’s engineers wrote up Camper, a system that drives multi-repository security campaigns by linking Jira, GitHub and automated remediators, with Temporal Workflows holding state so a campaign can pause, retry and resume without an operator. As of 29 September it had opened at least 150 public pull requests across 72 of the company’s repositories, 111 of which had merged. The tool is still private and early-stage, with open-sourcing contingent on operational maturity. It is a production account of workflow orchestration applied to long-running, human-reviewed engineering work rather than to an agent.</p>

<h3 id="weaviate-patches-a-high-severity-credential-leak-in-its-google-modules"><a href="https://weaviate.io/blog/weaviate-security-release-googlemodules-2026">Weaviate patches a high-severity credential leak in its Google modules</a></h3>

<p><strong>Weaviate</strong> · 1 Oct 2026 · <em>Vector stores</em></p>

<p>Weaviate 1.39.3 fixes a flaw where an unvalidated API endpoint setting in the text, multimodal and generative Google modules let a caller redirect outbound requests to an arbitrary host, which would receive the operator’s Google credentials as a bearer token. The issue is rated high severity (CVSS 7.1), could be triggered through collection configuration or GraphQL query parameters, and was not known to be exploited. Cloud and Marketplace customers were patched automatically; self-hosted users running earlier versions with those modules should upgrade.</p>]]></content><author><name></name></author><category term="data" /><category term="model-routing" /><category term="retrieval" /><category term="durable-execution" /><category term="vector-security" /><summary type="html"><![CDATA[LangChain’s model router cuts coding-agent cost by about two thirds with no measurable quality loss]]></summary></entry><entry><title type="html">Data and Orchestration Weekly — 27 September 2026</title><link href="https://pvelua.github.io/news/data/2026/09/27/data-weekly/" rel="alternate" type="text/html" title="Data and Orchestration Weekly — 27 September 2026" /><published>2026-09-27T00:00:00-07:00</published><updated>2026-09-27T00:00:00-07:00</updated><id>https://pvelua.github.io/news/data/2026/09/27/data-weekly</id><content type="html" xml:base="https://pvelua.github.io/news/data/2026/09/27/data-weekly/"><![CDATA[<h3 id="langchain-gives-managed-deep-agents-per-user-memory-and-credential-scoping"><a href="https://www.langchain.com/blog/langsmith-managed-deep-agents-whats-new">LangChain gives Managed Deep Agents per-user memory and credential scoping</a></h3>

<p><strong>LangChain</strong> · 24 Sep 2026 · <em>Agent frameworks</em></p>

<p>Managed Deep Agents 0.8 splits memory into two mounted paths, one shared across an agent’s users and one scoped per identity, with default access policies that block group channels and HTTP callers from reading user-level memory. Credential handling now distinguishes agent-owned keys shared across everyone from user-owned OAuth grants unique to each caller, resolved through a single connections call instead of custom auth code, and LangSmith ships 23 ready-made integrations including GitHub and Google Workspace. New HTTP channels let deployed agents receive webhook traffic from internal tools and customer portals, and Slack channels gained file transfer for logs and contracts. The release targets a problem specific to agents serving many people through one deployment: keeping each person’s context and credentials separate without a team rebuilding that plumbing itself.</p>

<h3 id="duckdb-becomes-a-built-in-adapter-in-dbts-new-rust-engine"><a href="https://duckdb.org/2026/09/22/dbt-fusion">DuckDB becomes a built-in adapter in dbt’s new Rust engine</a></h3>

<p><strong>DuckDB</strong> · 22 Sep 2026 · <em>Data engineering</em></p>

<p>dbt v2, the Rust-based rewrite of the transformation tool that reached general availability earlier in September, now ships DuckDB as a first-party adapter rather than requiring the community-maintained dbt-duckdb package, downloading and caching the driver automatically on first run. The pairing brings catalog-aware materializations backed by DuckLake and Iceberg REST catalogs, features the older Python adapter never had, plus a version of DuckDB pinned to the release so those capabilities stay stable across projects. Because dbt v2 already stores its own metadata as Parquet instead of large JSON manifests, teams can now build, test and publish transformation models entirely on a laptop without provisioning a warehouse. It is a small integration with an outsized reach, given how many data teams already default to dbt for SQL transformations.</p>

<h3 id="langsmiths-engine-now-proposes-and-tests-its-own-fixes-for-failing-agents"><a href="https://www.langchain.com/blog/langsmith-engine-v2-redteam">LangSmith’s Engine now proposes and tests its own fixes for failing agents</a></h3>

<p><strong>LangChain</strong> · 24 Sep 2026 · <em>Agent evaluation</em></p>

<p>LangSmith Engine v2 adds a private-beta red-teaming mode that probes deployed agents for hallucinations and system-prompt violations before they reach production, on top of its existing detection of inefficient tool-call trajectories and drifting error-rate, latency and cost trends. For Deployment customers, Engine now reproduces a reported failure, proposes a fix, tests it against the failing inputs, and iterates until it passes, surfacing the result as a one-click pull request rather than just a diagnosis. LangChain says Engine has analyzed over 70 million traces since its May launch, and reports it catches twice as many issues on its own IssueBench suite and produces fixes rated 25% more effective on Terminal-Bench than the prior version. It moves LangSmith from flagging agent problems toward closing the loop on them automatically.</p>

<h3 id="temporals-serverless-workers-can-now-run-inside-amazon-bedrock-agentcore"><a href="https://temporal.io/blog/amazon-bedrock-agentcore-with-temporal-serverless-workers">Temporal’s serverless workers can now run inside Amazon Bedrock AgentCore</a></h3>

<p><strong>Temporal</strong> · 21 Sep 2026 · <em>Agent infrastructure</em></p>

<p>Temporal released a prerelease integration letting its Serverless Workers use Amazon Bedrock AgentCore Runtime as a compute provider, so a Temporal Workflow can act as the durable control loop for an agent while AgentCore supplies the managed, scale-to-zero compute underneath. Model calls and tool operations run as Temporal activities with their own retry policies, and a published Python sample shows a TemporalAgent built on AWS’s Strands framework converting those operations into durable steps. Because workflow state lives in Temporal’s event history rather than on any single worker, capacity can scale with activity bursts and drop during idle periods without losing an agent’s progress mid-task. It answers a specific gap: durable-execution frameworks and serverless agent runtimes have mostly been evaluated separately, not run together.</p>

<h3 id="langsmith-adds-a-chronological-view-of-what-an-agent-actually-did"><a href="https://www.langchain.com/blog/langsmith-trajectories-tracing">LangSmith adds a chronological view of what an agent actually did</a></h3>

<p><strong>LangChain</strong> · 24 Sep 2026 · <em>Agent observability</em></p>

<p>LangSmith’s new Trajectories view collapses a session’s nested trace structure into a single chronological feed of messages and tool calls spanning a main agent and any subagents, working out of the box with LangChain, LangGraph, Deep Agents, OpenAI and Claude SDKs, and coding agents like Codex and Cursor. Teams can score a trajectory with an online evaluator, route it to a human annotation queue, or export it as a dataset for fine-tuning, directly from the same view. The framing is explicitly about debugging complexity that raw traces obscure: as the team put it, “a single session can span many user turns, tool calls, retries, and subagent handoffs.” It is a readability layer over trace data LangSmith already collected, aimed at making long agentic sessions inspectable without wading through nested runs.</p>]]></content><author><name></name></author><category term="data" /><category term="agent-frameworks" /><category term="data-engineering" /><category term="agent-evals" /><category term="durable-execution" /><category term="observability" /><summary type="html"><![CDATA[LangChain gives Managed Deep Agents per-user memory and credential scoping]]></summary></entry><entry><title type="html">Data and Orchestration Weekly — 20 September 2026</title><link href="https://pvelua.github.io/news/data/2026/09/20/data-weekly/" rel="alternate" type="text/html" title="Data and Orchestration Weekly — 20 September 2026" /><published>2026-09-20T00:00:00-07:00</published><updated>2026-09-20T00:00:00-07:00</updated><id>https://pvelua.github.io/news/data/2026/09/20/data-weekly</id><content type="html" xml:base="https://pvelua.github.io/news/data/2026/09/20/data-weekly/"><![CDATA[<h3 id="a-decision-model-not-a-language-model-beats-frontier-llms-as-an-agent-judge"><a href="https://www.langchain.com/blog/jev-agent-evals-langsmith">A decision model, not a language model, beats frontier LLMs as an agent judge</a></h3>

<p><strong>LangChain</strong> · 20 Sep 2026 · <em>Agent evals</em></p>

<p>LangSmith tested Jev, a non-generative “System One” model from TypeSafe AI that scores a given state and returns typed probabilities instead of producing text, as a judge for agent evaluations, benchmarking it against three LLM judges including Claude on a weather-agent test set. Jev matched human evaluators on 100% of pass/fail calls, against 99.8%, 96.4% and 80.0% for the LLM judges, and its quality-score variance ran 92 to 913 times lower. A full evaluation run cost $0.00035 and 0.44 seconds per call, versus $28.17 total for an equivalent run using Claude as judge. The team frames this as cheap enough to run against every production trace, though they caveat the results as early and narrow in scope.</p>

<h3 id="a-federated-multi-agent-system-lifts-chat-engagement-75-for-a-healthcare-navigator"><a href="https://www.langchain.com/blog/how-included-health-built-federated-agents-for-healthcare-navigation-with-deep-agents-and-langgraph">A federated multi-agent system lifts chat engagement 75% for a healthcare navigator</a></h3>

<p><strong>LangChain</strong> · 17 Sep 2026 · <em>Agent architecture</em></p>

<p>Included Health described “Dot,” a healthcare-navigation system built as a LangGraph “supergraph” router that hands conversations to specialized sub-workflows for urgent care, scheduling, referrals and behavioral health, with a filesystem-backed skills registry that loads abbreviated capability descriptions until an agent actually needs the full detail. Durable execution lets a human advocate pause and rejoin a conversation without losing context, and every exchange queues into LangSmith for clinician review. Since launch the company reports a 75% lift in chat engagement, care-routing agreement above its 95% target, and detection of over 99% of high-risk situations in clinical audits. It’s a concrete data point for federated, skills-registry-style multi-agent design in a regulated setting.</p>

<h3 id="weaviates-139-quantization-overhaul-cuts-vector-index-memory-45"><a href="https://weaviate.io/blog/4-bit-rotational-quantization">Weaviate’s 1.39 quantization overhaul cuts vector-index memory 45%</a></h3>

<p><strong>Weaviate</strong> · 17 Sep 2026 · <em>Vector search</em></p>

<p>Weaviate 1.39 adds 4-bit rotational quantization alongside its existing 8-bit and 1-bit options, using SIMD-accelerated Fast Walsh-Hadamard transforms to rotate vectors, a “centered” variant that subtracts the mean to preserve recall on anisotropic embeddings, and exact storage of the two largest-magnitude coordinates to bound outlier error. On a one-million-vector benchmark the new mode cut heap usage 45% versus 8-bit quantization, and its centered variant reached 96.8 recall@10 on a standard embedding set versus 93.5 uncentered, with recall holding steady out to 250 million vectors. It gives teams a middle point on the memory-versus-recall curve instead of a binary choice between 8-bit and 1-bit compression.</p>

<h3 id="pinecone-open-sources-a-framework-for-building-vector-quantizers-from-shared-primitives"><a href="https://www.pinecone.io/blog/vq-bench/">Pinecone open-sources a framework for building vector quantizers from shared primitives</a></h3>

<p><strong>Pinecone</strong> · 17 Sep 2026 · <em>Retrieval benchmarks</em></p>

<p>Pinecone released VQ-bench, an open benchmark framework, with an accompanying paper presented at VecDB@VLDB 2026, built on the observation that most published vector quantizers decompose into the same small set of reusable primitives. New quantization schemes can be assembled from those primitives in a few lines of code rather than implemented from scratch, and any combination gets automated evaluation across metrics. On a 1.3-million-vector benchmark, standard PQ and OPQ produced the lowest reconstruction error while EDEN matched E-RaBitQ’s recall at substantially faster encoding. It turns quantizer comparison, usually redone one-off by each research group, into a shared, extensible tool.</p>

<h3 id="duckdb-ships-an-official-claude-code-skill-for-sql-first-data-work"><a href="https://duckdb.org/2026/09/16/duckdb-skills.html">DuckDB ships an official Claude Code skill for SQL-first data work</a></h3>

<p><strong>DuckDB</strong> · 16 Sep 2026 · <em>Agent tooling</em></p>

<p>The DuckDB team published duckdb-skills, a Claude Code plugin that routes Claude’s data operations through the DuckDB CLI: reading files, running queries, converting formats, browsing cloud storage, working with spatial data and searching documentation, installed via <code class="language-plaintext highlighter-rouge">/plugin install duckdb-skills@claude-plugins-official</code>. Session state persists in a <code class="language-plaintext highlighter-rouge">state.sql</code> file of <code class="language-plaintext highlighter-rouge">ATTACH</code>, <code class="language-plaintext highlighter-rouge">USE</code> and <code class="language-plaintext highlighter-rouge">LOAD</code> statements, and failed queries are read back and retried automatically rather than surfaced raw to the user. It’s a concrete example of a data tool built to be operated by an agent rather than a person, with the plugin layer, not a model update, doing the work of making that reliable.</p>

<h3 id="langchains-deep-life-sci-wires-a-research-agent-into-600000-trials-and-41-million-papers"><a href="https://www.langchain.com/blog/agent-harness-life-sciences">LangChain’s Deep Life Sci wires a research agent into 600,000 trials and 41 million papers</a></h3>

<p><strong>LangChain</strong> · 17 Sep 2026 · <em>Agent frameworks</em></p>

<p>LangChain released Deep Life Sci, an open-source agent harness for clinical and lab scientists built on its Deep Agents framework, with access to over 600,000 ClinicalTrials.gov studies, 29 million PubMed abstracts and 12 million full-text PubMed Central articles, plus sandboxed sub-agents for code execution and support for lab file formats like FASTA and SMILES. It ships as a harness teams extend with their own internal data rather than a hosted product, with demonstrated workflows spanning RNA-seq analysis enriched with literature and clinical-trial comparison. The framing is explicit: reducing drug-development costs by betting an open harness beats a closed one for organizations that want to plug in proprietary data.</p>

<h3 id="a-single-qdrant-index-serves-multilingual-retrieval-without-a-translation-pipeline"><a href="https://qdrant.tech/blog/shift-multilingual-rag/">A single Qdrant index serves multilingual retrieval without a translation pipeline</a></h3>

<p><strong>Qdrant</strong> · 16 Sep 2026 · <em>Retrieval</em></p>

<p>Qdrant described SHIFT, a training-free technique that learns a per-language vector offset by averaging embedding differences between translation pairs, then applies it at indexing time to pull non-pivot-language documents toward a shared pivot space, with queries optionally shifted the same way. Tested on the XRAG dataset with a small multilingual embedding model, the correction lifted overall recall@10 from 0.195 to 0.297 and nearly tripled cross-language recall, at a modest cost to same-language recall, and held up under approximate HNSW search. It’s a cheap alternative to keeping translated copies of a corpus or moving to a much larger multilingual embedding model.</p>]]></content><author><name></name></author><category term="data" /><category term="agent-evals" /><category term="vector-quantization" /><category term="agent-frameworks" /><category term="retrieval" /><summary type="html"><![CDATA[A decision model, not a language model, beats frontier LLMs as an agent judge]]></summary></entry><entry><title type="html">Data and Orchestration Weekly — 14 September 2026</title><link href="https://pvelua.github.io/news/data/2026/09/14/data-weekly/" rel="alternate" type="text/html" title="Data and Orchestration Weekly — 14 September 2026" /><published>2026-09-14T00:00:00-07:00</published><updated>2026-09-14T00:00:00-07:00</updated><id>https://pvelua.github.io/news/data/2026/09/14/data-weekly</id><content type="html" xml:base="https://pvelua.github.io/news/data/2026/09/14/data-weekly/"><![CDATA[<h3 id="aws-lets-developers-write-straight-into-an-agents-long-term-memory"><a href="https://aws.amazon.com/about-aws/whats-new/2026/09/agentcore-memory-direct-ingest/">AWS lets developers write straight into an agent’s long-term memory</a></h3>

<p><strong>Amazon Web Services</strong> · 8 Sep 2026 · <em>Agent memory</em></p>

<p>Amazon Bedrock AgentCore Memory gained an IngestData API that accepts conversational or arbitrary JSON payloads and feeds them straight into an agent’s configured long-term memory strategies, skipping the earlier requirement to first log everything as a short-term memory event. Extracted records surface through the existing retrieval calls, with optional Kinesis notifications and a job-listing endpoint for reprocessing failed extractions. It decouples long-term memory from the live conversation log, so teams can backfill an agent’s memory from batch sources like activity logs rather than only from chat turns.</p>

<h3 id="weaviates-disk-based-vector-index-reaches-general-availability"><a href="https://weaviate.io/blog/hfresh">Weaviate’s disk-based vector index reaches general availability</a></h3>

<p><strong>Weaviate</strong> · 9 Sep 2026 · <em>Vector search</em></p>

<p>HFresh, Weaviate’s disk-based vector index, moved from technical preview to general availability in version 1.38, trading some query latency for a large cut in memory footprint on big collections. In the company’s own benchmark, a billion 256-dimension vectors needed about 239MB of heap under HFresh versus 6.67GB for a comparable uncompressed HNSW index, while quantized postings cut storage up to 32x against 32-bit floats. It targets teams whose vector collections have outgrown what they can justify keeping fully in memory.</p>

<h3 id="fastmcp-4-ships-alongside-a-field-guide-to-building-mcp-servers-that-dont-bloat-the-context-window"><a href="https://www.prefect.io/blog/is-your-mcp-server-actually-good">FastMCP 4 ships alongside a field guide to building MCP servers that don’t bloat the context window</a></h3>

<p><strong>Prefect</strong> · 9 Sep 2026 · <em>MCP / protocols</em></p>

<p>Prefect engineers marked the release of FastMCP 4, built on MCP’s newer stateless protocol, with a set of hard-won rules for MCP server design: start from zero tools rather than mapping every REST endpoint one-to-one, since that mapping is what produces token bloat, and route especially large APIs through a “code mode” of just two tools, search and execute. They also argue servers need CI evals run against real model calls — about $3 a run on Claude, by their estimate — and middleware-based auth before anything reaches production. The advice matters because it is aimed squarely at the gap between an MCP server that technically works and one an agent can use efficiently.</p>

<h3 id="langchain-gives-multi-agent-subagents-two-distinct-ways-to-inherit-context"><a href="https://www.langchain.com/blog/organizing-context-in-a-multi-agent-harness">LangChain gives multi-agent subagents two distinct ways to inherit context</a></h3>

<p><strong>LangChain</strong> · 8 Sep 2026 · <em>Agent orchestration</em></p>

<p>LangChain’s deepagents framework now exposes two context modes for subagents: “isolated,” which starts a subagent with only its task description in a fresh window, and “fork,” which hands it the supervisor’s full conversation history as a continuation. The post maps each mode to a role — isolated for independent verifiers and parallel researchers, fork for workers and memory-extraction agents that need the prior investigation — and notes forked subagents also benefit from prompt caching. It’s a concrete answer to a recurring multi-agent design question: how much of a supervisor’s context a delegated subagent should actually see.</p>

<h3 id="llamaindexs-two-pass-pattern-skips-expensive-ocr-on-most-pages"><a href="https://www.llamaindex.ai/blog/just-in-time-agentic-ocr">LlamaIndex’s two-pass pattern skips expensive OCR on most pages</a></h3>

<p><strong>LlamaIndex</strong> · 11 Sep 2026 · <em>Retrieval pipelines</em></p>

<p>For agents working across ad hoc data rooms of tens to hundreds of documents, LlamaIndex described running its free, layout-aware LiteParse tool first and reserving costlier VLM-based OCR only for pages LiteParse flags as complex. Across a test set of 84 SEC filings totaling over 12,000 pages, the first pass finished in 32 seconds and flagged roughly a fifth of pages for the expensive second pass. The company still recommends running full VLM OCR up front for large offline batch pipelines; this “retrieve first, then zoom in” pattern is aimed instead at smaller, ad hoc document sets.</p>

<h3 id="langchain-adds-per-caller-identity-to-its-managed-agent-credentials"><a href="https://www.langchain.com/blog/connections-managed-credentials-and-per-caller-identity-for-managed-deep-agents">LangChain adds per-caller identity to its managed agent credentials</a></h3>

<p><strong>LangChain</strong> · 9 Sep 2026 · <em>Agent orchestration</em></p>

<p>LangChain’s Managed Deep Agents gained Connections, a credential layer that replaces a single shared service-account key with per-caller identity, organized along two axes: agent-owned versus user-owned credentials, and static secrets versus OAuth grants. A deployed agent calls a single <code class="language-plaintext highlighter-rouge">connections.get()</code> to either resolve a cached per-user OAuth token or trigger a fresh authorization flow, without handling client registration itself. It closes a specific gap in agent deployments, where shared credentials show what an agent can do but not who asked it to do it.</p>]]></content><author><name></name></author><category term="data" /><category term="agent-memory" /><category term="vector-search" /><category term="mcp" /><category term="orchestration" /><summary type="html"><![CDATA[AWS lets developers write straight into an agent’s long-term memory]]></summary></entry></feed>