
Help DDEV Grow: Star Us on GitHub
If you use DDEV, a simple way to mention the project is to star the GitHub repository. Head over to github.com/ddev/ddev↗ and click Star. It takes five seconds, and it can help us - a star count is one of the things new users, sponsors, and and AI check before trusting an open source tool. If you're already a star, thank you!
DDEV IntelliJ/PhpStorm Plugin Lands in the DDEV Org
The DDEV Integration plugin for IntelliJ/PhpStorm↗, maintained by @AkibaAT, has been transferred into the ddev GitHub organization. This was on our 2026 plans list, and it's great to see it land. Awesome maintainer AkibaAT has kept the plugin in excellent shape, and this move gives it a permanent home alongside the rest of the DDEV ecosystem.
What's New on the Blog
Community Highlights
Knecht.works Ships Sandbox Rollback - Following up on last month's beta-tester call, the team at knecht.works has added sandbox rollback to their agency dashboard, letting each automated DDEV run boot into its own disposable environment. Read the update↗
TYPO3 Snapshot: Pull and Anonymize Production Data Locally - Ramon Herrmann released Snapshot, an open-source TYPO3 extension that pulls databases and fileadmin from live/staging into a local DDEV environment, with built-in anonymization for GDPR compliance. Read the announcement↗
Quick DDEV Previews: A Self-Hosted Preview Service - Matthias Andrasch built a proof-of-concept service that spins up DDEV preview environments from any branch of a connected GitHub repository, based on Samuel Reichör's technical work. Screencast: Using it on Hetzner VPS↗ • View the repo↗
Community Tutorials from Around the Web
- Migrating a Local WordPress Site to DDEV on Windows/WSL2 (Spanish) → Adam Martín walks through moving a WordPress project from Local to DDEV running inside WSL2, including database import, URL fixes, SSL certificates, and troubleshooting port conflicts. Read on dev.adammartin.es↗
- Installing DDEV on Linux (French) → An updated walkthrough covering Docker prerequisites and DDEV installation on Ubuntu/Debian, Fedora, and openSUSE, plus mkcert certificate setup. Read on kgaut.net↗
- Global Commands for Database Dumps and Remote Imports (French) →
ddev db-import and ddev db-export, a pair of global DDEV commands for restoring and exporting Drupal databases with drush cache-clear and login-link steps built-in, plus a follow-up set (db-prod-import, ssh-prod, and their preprod equivalents) for pulling a remote database in one step, packaged as the ddev-drupal-tools↗ add-on. Read the first post↗ • Read the follow-up↗
- DDEV + a-blogcms as a MAMP Alternative (Japanese) → An introduction to DDEV for a-blogcms developers used to MAMP, covering setup, useful commands, and Mailpit for email testing. Read on kazumich.com↗
- Running Drupal's GitLab CI Checks Locally → How Kalamuna's
ddev checks and ddev checks-fixes commands mirror the Drupal.org GitLab CI template, so code that passes locally passes in CI. Read on kalamuna.com↗
DDEV Training Starting Up Again This Fall
Live training is back for the fall, three sessions open to everybody.
Upcoming DDEV Live Contributor and User Training Sessions
Zoom Join Info:
Link: Join Zoom Meeting
Passcode: 12345
Events & Community
DrupalCamp Tokyo 2026 - ANNAI presented on AI-driven Drupal development and sustainable open-source CMS strategy, including using DDEV with git worktree to run parallel Drupal environments. Read the report↗ (Japanese) - for English coverage of git worktree with DDEV, see Contributor Training: git worktree for Multiple DDEV Projects and Using git worktree with TYPO3.
Governance
- The next DDEV advisory group meeting, open to everybody, is September 2, 2026 at 8:00 AM US Mountain / 10:00 AM US Eastern / 16:00 CEST. Add to Google Calendar • See the agenda. We love to hear from our community!
Sponsorship Update
A steady month - thank you to everyone who contributes!
July 2026: ~$9,931/month (82.8% of goal)
August 2026: ~$10,038/month (83.7% of goal)
If DDEV has helped your team, consider sponsoring. → Become a sponsor↗
Contact us to discuss sponsorship options that work for your organization.
Stay in the Loop-Follow Us and Join the Conversation
Compiled and edited with assistance from Claude Code.
27 Aug 2026 12:00am GMT
I added a new feature to my blog: a list of related posts at the bottom of each post. I implemented it using embeddings, and this note documents how.
I looked at how other content management systems identify related posts: most use shared tags, backlinks, manual curation, or embeddings. I chose embeddings, which compare the meaning of each post, because they can uncover connections without shared tags, existing links, or manual curation.
Embeddings turn meaning into numbers
An embedding model reads text and returns a vector: a long list of numbers. The model I use, bge-base-en-v1.5 from the Beijing Academy of Artificial Intelligence (BAAI), returns 768 numbers for each post. I started with a smaller model that returns 384 numbers and moved up because the matches were better. BAAI's own benchmarks point the same way, though the gap is modest.
You can think of those 768 numbers as coordinates in a high-dimensional meaning space, where each dimension captures some pattern the model learned from text. For one of my posts, the first handful of those coordinates looks something like this:
[ 0.021, -0.045, 0.038, -0.012, 0.007, ..., 0.019 ] (768 numbers total)
Conceptually, it is a bit like tagging each blog post with hundreds of auto-generated tags, except that these tags are unnamed (they are just numbers) and distributed (meaning is spread across all of them). Together, the 768 numbers place the post near other posts with similar meaning.
This is what lets two posts match even when they use different words. During training, the model learns that certain words and phrases appear in similar contexts or play similar roles, so it places them near each other in the space. It does not need "car" and "automobile" to share any letters to learn that they are used in related ways.
Raw cosine similarity makes everything look related
Once every post has an embedding vector, the next question is how to compare them. This is where I had to dust off a little math. Fortunately, it turned out to be mostly high-school math: averages, angles, and multiplication.
The standard way to compare two vectors is cosine similarity. Imagine each vector as an arrow pointing away from the origin. Cosine similarity measures the angle between two of these arrows and then takes the cosine of that angle, which is where the name comes from.
Two arrows pointing almost the same way sit at a small angle, and the cosine of a small angle is close to 1, so the posts are related. As the arrows spread apart, the cosine falls: at a right angle it is 0, and for arrows pointing in opposite directions it drops to -1, so unrelated posts score closer to 0 or even negative.
In practice, these raw cosine values can be misleading, because embedding models rarely spread their vectors evenly in every direction. They tend to pack most vectors into a narrow cone, a property called anisotropy, so the scores cluster in a high, narrow band. On my blog, the raw cosine similarity between two randomly chosen posts is almost always between 0.5 and 0.75, with a median of 0.64.
The practical effect is that almost any two posts look somewhat similar. An old post about the founding of Acquia shows the problem. It covers a lot of ground: Drupal, my PhD, Red Hat and IBM backing Linux, venture capital, and personal reflection. Because it touches so many subjects, its vector sits close to the average of all my posts, and it scored high against almost the entire archive. Its best match scored 0.876, and its hundredth best still scored 0.770.
Mean-centering reveals what makes each post distinct
Anisotropy has several known fixes, from lightest to heaviest. The lightest is mean-centering, which is what I use and what the rest of this section explains.
All-but-the-top removes the average and the next few strongest directions. Whitening stretches the space so every direction carries equal weight (the name comes from white noise). I have not tried these others. Mean-centering is one subtraction per vector with no matrix algebra, which keeps the code plain PHP, and it was enough.
You compute the average vector across all posts and subtract it from every post's vector. Subtracting the average vector from each post removes what all posts have in common, so what remains is what makes each post distinct. That average points down the middle of the cone, the direction my whole blog tends to lean.
A modern model like bge-base-en-v1.5 already suffers less from anisotropy than older or simpler encoders: it is trained with contrastive learning, which pushes unrelated texts apart, and version 1.5 was tuned specifically to spread out its similarity scores. On my corpus, centering still made the scores much more useful.
An example might help. Imagine three posts with only two numbers each instead of 768:
A = (0.90, 0.10)
B = (0.85, 0.80)
C = (0.80, 0.75)
At first glance, all three posts look somewhat similar. In every post the first number is high and close to the others (0.90, 0.85 and 0.80), so it dominates the comparison. But a number that barely changes from post to post tells you little about how they differ, so that first number is not very useful.
The average (mean) of the three vectors is:
mean = (0.85, 0.55)
Now subtract that average from each post:
A = ( 0.05, -0.45)
B = ( 0.00, 0.25)
C = (-0.05, 0.20)
Now the picture is clearer. B and C both have a positive second number, so they point in roughly the same direction; A's second number is negative, so it points somewhere else.
Before centering, everything looked similar. After centering, the comparison focuses on what is different from the average.
Normalization reduces comparison to a dot product
After centering, each vector has a length as well as a direction. Length says how far a post sits from the average, and direction says in what way it differs.
I want to rank posts by what they are about, not by how unusual they are, so only the direction matters. Hence, we normalize each vector by dividing it by its own length, which scales it to length 1 and moves it onto the unit circle (or, in 768 dimensions, the unit sphere), leaving only its direction.
It also makes the comparison cheaper. Cosine similarity is normally the dot product divided by the product of the two vectors' lengths. If both vectors have length 1, that denominator is 1 × 1 = 1, so the expression reduces to the dot product alone: multiply the two lists number by number, then add the results.
Using the same example, the centered vectors for B and C are:
B = ( 0.00, 0.25)
C = (-0.05, 0.20)
First, normalize each vector to length 1. A vector's length is the square root of the sum of its squared numbers (good old Pythagoras, only with more numbers). B has length √(0.00² + 0.25²) = 0.25, while C has length √((-0.05)² + 0.20²) ≈ 0.206, so dividing each vector by its own length gives:
B ≈ ( 0.00, 1.00)
C ≈ (-0.24, 0.97)
Then take the dot product:
(0.00 × -0.24) + (1.00 × 0.97) = 0.97
That is a strong match: the closer the score is to 1, the more the two posts point in the same direction. B and C are nearly aligned.
A, after normalization, points mostly downward. Next to B:
A ≈ ( 0.11, -0.99)
B ≈ ( 0.00, 1.00)
Multiplying them the same way:
(0.11 × 0.00) + (-0.99 × 1.00) = -0.99
That is not a match at all.
The PHP code is shorter than the explanation
The production code does the same arithmetic, just with 768 numbers per post instead of two:
public static function center(array $raw): array {
if ($raw === []) {
return [];
}
$mean = array_fill(0, count(reset($raw)), 0.0);
foreach ($raw as $vector) {
foreach ($vector as $i => $value) {
$mean[$i] += $value;
}
}
$count = count($raw);
foreach ($mean as $i => $sum) {
$mean[$i] = $sum / $count;
}
$centered = [];
foreach ($raw as $nid => $vector) {
$norm = 0.0;
foreach ($vector as $i => $value) {
$vector[$i] = $value - $mean[$i];
$norm += $vector[$i] * $vector[$i];
}
// A vector sitting exactly on the mean centers to zero; fall back to 1.0
// so the division below never hits a zero norm.
$norm = sqrt($norm) ?: 1.0;
foreach ($vector as $i => $value) {
$vector[$i] = $value / $norm;
}
$centered[$nid] = $vector;
}
return $centered;
}
public static function topMatches(array $source, array $pool, int $self): array {
$scores = [];
foreach ($pool as $nid => $vector) {
if ($nid === $self) {
continue;
}
$similarity = 0.0;
foreach ($source as $i => $value) {
$similarity += $value * $vector[$i];
}
$scores[$nid] = $similarity;
}
arsort($scores);
return array_keys(array_slice($scores, 0, 3, TRUE));
}
While my explanation was long, both PHP methods are relatively short. In center(), each vector has the corpus mean subtracted, then is divided by its own length. In topMatches(), I calculate the cosine similarity between one post and every other post, then keep the three highest.
You might expect a vector database to replace all of this. It would replace some of it: storing a vector and asking for the closest three would remove topMatches(), but it would not remove center(). Centering is optional, but it meaningfully improved my results.
A vector database likely makes centering harder. Today I store raw vectors and subtract the average when I compare them, so a new post does not change anything I have stored. A vector database would search what I stored, so the subtraction would have to happen before storing. I'd have to update all stored vectors for every new post or every edit, which feels more complex. Maybe vector databases have a good answer for that; I have not looked.
One-time embeddings, occasional ranking
You might wonder how expensive it is to generate these embeddings and compare all these vectors. It turns out to be fast and cheap.
There are two kinds of work, and they happen at different times. Generating an embedding calls an AI model, but happens only once after a post is created or edited. Ranking uses ordinary PHP arithmetic and happens occasionally, when Drupal rebuilds a page's cached related-post list.
I run the model on Cloudflare Workers AI. To generate an embedding, my server makes an HTTPS call that passes the post's text to Cloudflare, which runs the model and returns the 768-number vector. That round trip takes about 250ms. It happens on the first view after a post is created or edited, and the vector is then cached. The model is deterministic, so the same text always produces the same 768 numbers.
Cloudflare bills Workers AI usage in units it calls Neurons and includes 10,000 free each day. Embedding my full archive of roughly 1,500 posts used roughly 4,000 Neurons, and a new post costs about three. Embedding my blog is basically free.
Calculating the related posts never calls the AI model. It all happens in Drupal, my website's content management system. When Drupal needs to build one of the related posts lists, it loads all the stored vectors, centers them, and scores the current post against all the others: roughly 1,500 dot products, each over 768 numbers. This takes around 250ms on my site. After a list has been built, it is cached.
In other words, my website never loads model weights; it just stores the 768 numbers that come back. The machine-learning compute lives at Cloudflare's edge, and my server stays a plain PHP application. None of this needs a vector database or a machine-learning framework: one HTTP call generates the embedding, a key-value store caches it, and a few dozen lines of arithmetic choose the related posts.
Tags are too blunt, backlinks only capture the links I remembered to make, and manual curation does not scale. All three need me to notice the connection first. Using embeddings might sound a bit scary, but they turned out to be easy to implement, fully automated, and able to surface posts I would never have thought to link.
26 Aug 2026 8:40am GMT