<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>VC Deal Flow Signal: Blog</title>
    <link>https://signals.gitdealflow.com/blog</link>
    <description>Insights on using GitHub engineering signals for startup investing. Practical guides for VCs and angel investors.</description>
    <language>en</language>
    <atom:link href="https://signals.gitdealflow.com/feed.xml" rel="self" type="application/rss+xml"/>
    <atom:link href="https://pubsubhubbub.appspot.com/" rel="hub"/>
    <atom:link href="https://pubsubhubbub.superfeedr.com/" rel="hub"/>
    <atom:link href="https://signals.gitdealflow.com/atom.xml" rel="alternate" type="application/atom+xml"/>
    <atom:link href="https://signals.gitdealflow.com/feed.json" rel="alternate" type="application/json"/>
    <lastBuildDate>Tue, 25 Aug 2026 13:58:50 GMT</lastBuildDate>
    <item>
      <title><![CDATA[Venture Scouting: How Scouts and Angels Source Deals Before the Databases]]></title>
      <link>https://signals.gitdealflow.com/blog/venture-scouting-guide</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/venture-scouting-guide</guid>
      <description><![CDATA[How venture scouts and angels source startup deals ahead of Crunchbase and the funding databases, from scout programs to referral networks and leading-indicator data.]]></description>
      <content:encoded><![CDATA[<p>The entire scouting edge is timing. A database tells you what already happened; a scout earns their place by seeing it before it happens. This is the sourcing discipline, not the capital, and it is learnable.</p>
<h2>What a Scout Actually Does</h2>
<p>A scout sources deals, vets them against a fund's thesis, and refers the ones that clear the bar, in exchange for carry or a referral fee. The job is two skills: reach into places the fund is not looking, and judgment about which companies are worth a real look.</p>
<h2>Scout Programs</h2>
<p>Scout programs formalize the arrangement. A fund gives a scout a small allocation, and the scout deploys it on deals they source. The scout's edge is not capital; it is access to a network, a community, or a signal the fund does not already have.</p>
<h2>The Sourcing Stack</h2>
<p>The best scouts run several sourcing channels at once. Referral networks bring warm intros. Communities and forums surface companies before they raise. Public data, including engineering activity, surfaces companies whose shipping pace is accelerating before the round is announced <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Leading Indicators vs. Lagging Databases</h2>
<p>Databases are lagging indicators by design: they record rounds after they are announced. The scouting opportunity is the leading indicator, the 3 to 6 weeks before the announcement when a company's engineering activity is already accelerating. The backtest against 219 fundraises found the signal precedes Series A by 21 to 47 days <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>Where Public Engineering Data Fits</h2>
<p>A weekly-updated public panel of 350+ startup GitHub orgs gives a scout a discovery surface that does not depend on their referral graph. The free MCP server exposes six read-only tools for trending startups, sector search, startup lookup, signal summary, scout receipts, and methodology, usable inside any agent runtime <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<p>Scouting is not a network you are born into. It is a system for seeing deals before the databases do, and public engineering data is one of the cheapest inputs to that system.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Pre-Seed Scouting With GitHub Signals: Finding Startups Before Their First Announcement]]></title>
      <link>https://signals.gitdealflow.com/blog/pre-seed-scouting-with-github-signals</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/pre-seed-scouting-with-github-signals</guid>
      <description><![CDATA[How to scout pre-seed startups using public GitHub signals, including the thin-footprint caveat that makes pre-seed a verification game rather than a discovery game.]]></description>
      <content:encoded><![CDATA[<p>Pre-seed is where the scouting model breaks down first, for a simple reason: there is almost nothing public to read. The team has not announced, the press has not written, and the code footprint is thin.</p>
<p>Understanding that limitation is what separates a useful pre-seed workflow from a wasted one.</p>
<h2>Why Pre-Seed Is the Hardest Stage</h2>
<p>At pre-seed, a startup is two to four engineers, often working partly in private repositories. The public footprint is thin by definition, and the engineering signal that is discovery-grade at Series A is too weak to surface a company cold at pre-seed <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>What Public GitHub Shows at Pre-Seed</h2>
<p>What it does show is whether a team is real and whether it is shipping. Consistency matters more than volume: a founder who commits regularly at low absolute counts is a stronger signal than a burst followed by silence. Contributor breadth tells you whether it is a solo founder or a team, and repository activity tells you whether the product story matches the code <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<h2>Verification, Not Discovery</h2>
<p>The honest framing is that public GitHub activity at pre-seed is a verification tool, not a discovery tool. Use it to confirm that a founder you met through a community, a referral, or an accelerator is actually building, not to find the company in the first place <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>A Pre-Seed Scouting Workflow</h2>
<p>Source through the channels that do find pre-seed companies: communities, referrals, accelerators, and founder circles. When a founder surfaces, run a free public pass on their GitHub before you spend a meeting. Check that the team is real, the product story matches the code, and the shipping is consistent.</p>
<h2>The Early-Warning Use</h2>
<p>There is one case where the pre-seed signal is worth tracking proactively: acceleration. When a pre-seed team you already know begins to accelerate, they are usually deploying capital or approaching a round, which is exactly the moment a scout wants to be in the conversation <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>Pre-seed scouting is hard because the record is thin. But the cheapest check in the stack, a public GitHub pass, is also the one that keeps you from wasting a meeting on a team that is not actually building.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Engineering Velocity Benchmarks: What Fast Looks Like on GitHub, by Stage and Sector]]></title>
      <link>https://signals.gitdealflow.com/blog/engineering-velocity-benchmarks-by-stage</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/engineering-velocity-benchmarks-by-stage</guid>
      <description><![CDATA[Reference numbers for reading startup engineering velocity on GitHub: how sectors differ, what the acceleration thresholds are, and how to benchmark a team against its own baseline.]]></description>
      <content:encoded><![CDATA[<p>Benchmarks exist for one reason: to tell the difference between a team that looks busy and a team that is actually accelerating. Without a reference point, a commit count is just a number.</p>
<p>Here are the reference points the weekly panel actually uses, and where each one is and is not useful.</p>
<h2>The Panel</h2>
<p>GitDealFlow tracks public GitHub activity across a weekly panel of 350+ startup orgs in 15 sectors, from web3 at 42 startups to agtech at 11 in the current Q3 2026 snapshot <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>. The composition reflects where public engineering activity is most observable, not where the best companies are.</p>
<h2>The Acceleration Screen</h2>
<p>The weekly Acceleration Watch cohort uses a transparent threshold: at least 40 commits per 14-day window and at least 5 contributors, excluding percentage jumps off a near-zero baseline. The threshold is a screen against low-base artifacts, not a definition of fast in absolute terms.</p>
<h2>Sector Differences</h2>
<p>Baseline velocity differs by sector. Enterprise SaaS teams commit 35 to 60 percent less than AI-tools teams at the equivalent stage, because compliance and change-control processes throttle deployment frequency in a way public commit data reflects directly <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup>. This is why sector-stratified ranking beats a single global percentile.</p>
<h2>Stage Differences</h2>
<p>Velocity and contributor behavior also shift with stage. Pre-seed teams are thin and lumpy. Seed and Series A teams show the sharpest acceleration as they deploy capital into headcount. Growth teams show contributor expansion more than velocity spikes. The stage tells you which dimension to weight.</p>
<h2>The Best Benchmark Is a Team Against Itself</h2>
<p>The single most useful benchmark is a team against its own prior pace. A company accelerating from 18 to 36 commits per 14 days is a top-decile move inside enterprise SaaS even though the absolute number looks small. Trajectory, not level, is the signal <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>Benchmarks are reference points, not verdicts. They tell you what fast looks like so you can recognize acceleration when you see it.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Commit Velocity Benchmark Numbers: The Thresholds Behind a Startup Acceleration Signal]]></title>
      <link>https://signals.gitdealflow.com/blog/commit-velocity-benchmark-numbers</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/commit-velocity-benchmark-numbers</guid>
      <description><![CDATA[The specific commit velocity numbers behind the GitDealFlow acceleration signal: observation windows, contributor thresholds, and how to read velocity change.]]></description>
      <content:encoded><![CDATA[<p>Commit velocity is the raw throughput metric behind the acceleration signal, but a raw number without windows and thresholds is noise. Here are the specific numbers that make it interpretable.</p>
<h2>What Commit Velocity Counts</h2>
<p>Commit velocity is the count of commits pushed to a startup's public repositories within a rolling window, read through the public GitHub statistics API <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>. It is a proxy for shipping pace. It does not measure code quality, and it was never meant to.</p>
<h2>The 14-Day and 28-Day Windows</h2>
<p>The default read is a 14-day window. But teams that work in two-week sprints produce naturally lumpy distributions, a heavy push followed by quiet, and a 14-day window can catch half a sprint and misread the quiet half as deceleration. The 28-day window smooths that artifact and gives a cleaner read of whether overall pace is changing <sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup>.</p>
<h2>The Contributor Floor</h2>
<p>A percentage change only means something above an absolute floor. A jump from 1 to 4 commits is a 300 percent change that means nothing. A minimum contributor count, five in the weekly screen, suppresses the single-author artifacts that would otherwise dominate the signal <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Reading Velocity Change</h2>
<p>Read the change against the team's own baseline, not against the whole panel. A team moving from 18 to 36 commits per 14 days is a top-decile move inside enterprise SaaS even though the absolute number looks modest. And confirm a move across both windows: a one-window spike is worth a look, a two-window trend is worth a call <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>Why the Numbers Are Thresholds, Not Rankings</h2>
<p>None of these numbers is a grade. The commit floor, the contributor floor, and the windows are filters that keep the signal honest. They exist so that when a startup shows up as accelerating, the acceleration is real rather than an artifact of a low baseline <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<p>Commit velocity is the input. Contributor growth and repository expansion are the confirmation. The composite of the three is what actually predicts a fundraise.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[How to Evaluate Startup Founders: Signals That Predict Execution]]></title>
      <link>https://signals.gitdealflow.com/blog/how-to-evaluate-startup-founders</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/how-to-evaluate-startup-founders</guid>
      <description><![CDATA[A framework for evaluating startup founders before writing a check: execution signals, technical depth, team-building, and the red flags that matter at seed.]]></description>
      <content:encoded><![CDATA[<p>At seed stage, the founder is the investment. The product will pivot, the market will shift, but the team's ability to execute is the constant. Which makes founder evaluation most of the job.</p>
<h2>Why Founders Matter Most at Seed</h2>
<p>A seed check is a bet on a team before the product and market have proven themselves. Everything else on the scorecard, market, product, traction, is a bet on the founder's ability to find them. Get the team wrong and nothing else matters.</p>
<h2>Execution Signals</h2>
<p>The best predictor of whether a founder can deliver is whether they are delivering right now. That is readable in their public shipping record: the commits they push, the repositories they build, and the pace at which they do it. It is the least gameable founder check available <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Technical Depth</h2>
<p>For technical founders, public code is a direct record of capability. It shows what they actually build, at what pace, and with how many collaborators. It is also the cheapest check: reading a founder's public GitHub costs nothing and cannot be rehearsed.</p>
<h2>Team-Building Signals</h2>
<p>A founder who can recruit is a founder who can scale. Watch contributor growth on their projects. A founder whose team is expanding is building something other people want to work on, which is itself a signal about the company.</p>
<h2>The Red Flags</h2>
<p>A few signals override everything else. A pitch that contradicts the public record. A shipping pace that does not match the claimed traction. A founder who cannot name the specific problem their product solves. Any one of these is worth walking away over.</p>
<h2>Reading Founders in Public</h2>
<p>A weekly-updated panel of 350+ startup orgs gives you a surface for spotting founders who ship, independently of their network <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>. It does not replace the conversation, but it decides who earns one <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<p>Founder evaluation is judgment, but the inputs to that judgment are increasingly public. Read the shipping record first, then take the meeting.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Assessing Technical Founders From Their GitHub: Shipping Discipline as a Quality Signal]]></title>
      <link>https://signals.gitdealflow.com/blog/technical-founder-assessment-github</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/technical-founder-assessment-github</guid>
      <description><![CDATA[How to assess a technical founder's capability from their GitHub activity: shipping discipline, contributor patterns, and the Scout Score that grades a user's star history.]]></description>
      <content:encoded><![CDATA[<p>A technical founder's GitHub is the most honest resume they have. It is a running record of what they build, at what pace, and with whom, generated daily without any intent to impress. Reading it well is a founder-evaluation superpower.</p>
<h2>Why GitHub Is the Honest Resume</h2>
<p>LinkedIn can be curated. A pitch deck can be rehearsed. But a commit history is a record of actual work, produced continuously, and it cannot be faked for an interview. That makes it the single least gameable signal in founder evaluation <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<h2>Shipping Discipline</h2>
<p>The first thing to read is discipline: consistent commits over time, not a burst followed by silence. A founder who ships steadily at low absolute volume is a stronger signal than a founder who ships in a burst and then goes quiet. Consistency is execution made visible <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>Contributor Patterns</h2>
<p>Next, read who they build with. Does the founder build alone, or do other people choose to work with them? Contributor growth on their projects is a signal about both the founder's recruiting and the quality of the work, because people do not contribute to projects that are not worth their time.</p>
<h2>Scout Score: Grading a Founder's Eye</h2>
<p>There is a second dimension the GitHub record reveals: judgment. The Scout Score grades a GitHub user's star history against roughly 75 validated unicorns, measuring whether they have been watching breakout companies before they were obvious. A founder who consistently spotted winners early has demonstrated the judgment that seed investing depends on <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Honest Limits</h2>
<p>GitHub tells you what a founder builds and whether they can spot a winner. It does not tell you how they handle pressure, whether they are coachable, or how they treat their team. Those still require a conversation.</p>
<h2>The Workflow</h2>
<p>Check the trailing activity for shipping discipline, read the contributor patterns for team-building, look at the specific repositories to confirm the product story, and run the Scout Score for judgment. Then take the meeting with a founder you already have reasons to believe in <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<p>Founder evaluation is a judgment call, but it is a judgment call you can now make with better inputs. The public record is the cheapest, most honest input of all.</p>]]></content:encoded>
      <pubDate>Fri, 14 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Deal Flow Management for Early-Stage Investors: Capture, Triage, Score, Prioritize]]></title>
      <link>https://signals.gitdealflow.com/blog/deal-flow-management-for-early-stage-investors</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/deal-flow-management-for-early-stage-investors</guid>
      <description><![CDATA[A practical deal flow management system for angels and early-stage VCs: how to capture, triage, score, and prioritize inbound startups without a full-time analyst.]]></description>
      <content:encoded><![CDATA[<p>Most early-stage investors do not have a deal flow problem. They have a deal flow management problem. The same companies keep arriving, the same ones keep getting a first look, and the best ones keep slipping because there is no system deciding where attention goes.</p>
<p>Here is a system that works with a spreadsheet and one hour a week.</p>
<h2>Capture: Every Deal Lands in One Place</h2>
<p>If a deal lives only in your inbox or your memory, it is not in your pipeline. Capture means one list where every inbound opportunity lands, tagged with source, sector, stage, and the date it arrived. The source tag matters more than it seems: it tells you later which channels actually produce the deals you close.</p>
<h2>Triage: One Question, One Minute</h2>
<p>Triage is the filter that keeps the rest of the pipeline honest. For each new deal, answer a single question in under a minute: does this clear the bar for a real look, yes or no. Everything that clears moves forward. Everything else is archived with a one-line reason, because a written reason is what lets you calibrate the bar later.</p>
<h2>Score: Objective Before Subjective</h2>
<p>Score the survivors on what is public before you spend meeting time. Market, product, traction, and shipping trajectory are all readable without a call. Public GitHub activity gives you the trajectory dimension for free: a weekly-updated read of commit velocity and contributor growth across 350+ startup orgs <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Prioritize: The Weekly Sort</h2>
<p>Once a week, sort the scored pipeline and decide where your next block of time goes. The review is not optional; it is the rhythm that keeps a pipeline from rotting into a list of companies you once looked at.</p>
<h2>Why Public Signals Fit the Pipeline</h2>
<p>Sourcing is not the same as managing, but a public, machine-readable panel changes the math on both. It lets you discover companies outside your referral graph and check shipping trajectory before you commit a meeting. It does not replace your judgment; it gives your triage stage better inputs <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>A pipeline with a capture list, a one-minute triage, a public-first score, and a weekly sort will outperform a larger, messier pipeline every time. The system is the edge, not the volume.</p>]]></content:encoded>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[A Deal Flow Scoring Framework: Rank Inbound Startups Without a Full Partner Meeting]]></title>
      <link>https://signals.gitdealflow.com/blog/deal-flow-scoring-framework</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/deal-flow-scoring-framework</guid>
      <description><![CDATA[A four-factor deal flow scoring framework for early-stage investors, with a fifth engineering-velocity factor from public GitHub data, plus calibration guidance.]]></description>
      <content:encoded><![CDATA[<p>A scorecard is not a substitute for judgment. It is a way to make judgment explicit, comparable, and reviewable, so that the decision you make on deal forty is as disciplined as the one you made on deal one.</p>
<p>Here is a framework that is fast enough to use on every inbound deal and structured enough to calibrate.</p>
<h2>Score Before You Meet</h2>
<p>The point of scoring is to rank before you spend the expensive meeting. Score on what is public, then decide who earns the call. A 0 to 10 scale on each factor keeps the pass fast enough that you will actually do it.</p>
<h2>The Four Core Factors</h2>
<p><strong>Team.:</strong> Do the founders have the right capability and track record for this specific problem. At seed, this is the highest-weight factor.</p>
<p><strong>Market.:</strong> Is the market real, large, and growing. A great team in a bad market is still a bad investment.</p>
<p><strong>Product.:</strong> Does the product exist and solve the stated problem. A live product beats a deck, every time.</p>
<p><strong>Traction.:</strong> Is there evidence of real usage, revenue, or engagement. Traction weight rises with stage.</p>
<h2>The Fifth Factor: Engineering Velocity</h2>
<p>Public GitHub activity adds a leading-indicator dimension the other four do not cover. It reads whether the team is shipping faster or slower, independently of what the deck claims. The backtest against 219 fundraises found a 3.4x lift in a composite commit-velocity and contributor signal preceding Series A, with a 21 to 47 day lead <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>Score it on trajectory: accelerating high, flat middle, decelerating low. It is the only factor on the card that predicts rather than reports.</p>
<h2>Setting Weights</h2>
<p>Write the weights down. At pre-seed, team might be 40 percent with velocity and product splitting the rest. At Series A, traction and market rise. The weights are a policy; the discipline of writing them is what makes the scorecard useful later.</p>
<h2>Calibrate or Discard</h2>
<p>A scorecard that is never checked against outcomes is superstition. Go back after six months and test whether high scores predicted raises and performance. Adjust the weights where the card was wrong.</p>
<p>The goal is not a perfect number. It is a decision process you can defend, review, and improve <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>]]></content:encoded>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[The 10 Best Chrome Extensions for VC Deal Flow (2026)]]></title>
      <link>https://signals.gitdealflow.com/blog/best-chrome-extensions-vc-deal-flow-2026</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/best-chrome-extensions-vc-deal-flow-2026</guid>
      <description><![CDATA[The 10 Chrome extensions venture investors actually use in 2026 to source deals, research startups, and move faster, including two purpose-built for GitHub engineering signal. Real pricing, honest disclosure, install links.]]></description>
      <content:encoded><![CDATA[<p>If you source venture deals, your browser is your office. The right Chrome extensions turn Crunchbase, LinkedIn, and GitHub from passive databases into an active deal-flow engine, surfacing signals, enriching contacts, and cutting hours off diligence per company.</p>
<p>This is the list we wish we'd had when we started. It includes two extensions we built ourselves (we'll be upfront about that) and eight others that earn their place in a working VC stack. Every entry has real pricing and a clear &quot;why an investor needs it&quot;, no fluff, no paid placements.</p>
<h2>1. VC Deal Flow Signal, GitHub Startup Signals (free)</h2>
<p><strong>What it does:</strong> Overlays a live engineering-signal badge on Crunchbase and Wellfound company profiles. Open a company page and the badge shows whether the startup's engineering is accelerating, steady, or decelerating against its own baseline. Hover for the underlying metrics: 14-day commit velocity, velocity change versus the prior period, contributor count and growth, and the signal type (hiring burst, infrastructure buildout, framework migration, deploy-frequency spike) <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<p><strong>Pricing:</strong> Free in perpetuity. No account, no API key, no tracking. Deeper signal lives in the paid tiers (€7 First Look → €49/month Dashboard → €197/month Insider Circle), each with a 30-day guarantee.</p>
<p><strong>Why investors need it:</strong> Crunchbase tells you what already happened, the last round, the announced valuation. It doesn't tell you whether the engineering team is accelerating right now. This badge layers a leading indicator on the lagging database you already read: public GitHub activity across a 350+ startup panel <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>, where sustained acceleration typically precedes a fundraise announcement by three to six weeks <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>. If you research deals on Crunchbase, install this first.</p>
<h2>2. VC GitHub Lookup, Startup Signals on Hover (free)</h2>
<p><strong>What it does:</strong> Puts the same signal on GitHub itself. Hover any GitHub repo or org and a chip appears with commit velocity, contributor growth, and breakout status; direct visits to an org or repo page get the chip too, and the toolbar popup runs a manual lookup against any GitHub URL <sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup>.</p>
<p><strong>Pricing:</strong> Free. Same paid ladder as #1.</p>
<p><strong>Why investors need it:</strong> Half of early-stage technical diligence ends with &quot;let me check their GitHub.&quot; Native GitHub shows stars and a repo list; it does not tell you whether the team is ramping or stalling against its own baseline. This turns a 15-minute scroll through commit history into a two-second read, on the same weekly-refreshed panel as the Crunchbase badge.</p>
<p><strong>Disclosure:</strong> #1 and #2 are GitDealFlow's own extensions. We built them because nothing else did this. The remaining eight are third-party tools that round out the stack.</p>
<h2>3. Affinity (paid, from ~$100/user/mo)</h2>
<p><strong>What it does:</strong> A relationship-intelligence CRM whose extension captures every email, meeting, and contact interaction automatically and links them to deals in your pipeline. No manual data entry.</p>
<p><strong>Why investors need it:</strong> The most common reason deals die is that nobody followed up. Affinity silently logs who you talked to, when, and about what, so when a founder resurfaces six months later, the full history is one click away. It's the CRM many growth-stage firms standardize on; pricing is custom and starts around $100 per user per month.</p>
<h2>4. Clearbit Connect (free tier + paid)</h2>
<p><strong>What it does:</strong> Reveals email address, role, and company details for people at the company you're viewing, from a browser sidebar. The free tier carries a limited monthly lookup budget.</p>
<p><strong>Why investors need it:</strong> When a startup looks interesting, the next step is reaching the founder or VP Engineering, not the info@ black hole. Clearbit surfaces verified addresses enriched with title, seniority, and headcount so the first touch lands with the right person.</p>
<h2>5. Hunter (free tier + paid)</h2>
<p><strong>What it does:</strong> Shows the email pattern a company uses (first.last@company.com and variants) and verifies deliverability before you send. The free tier includes a small monthly search budget; paid plans start around $34/month.</p>
<p><strong>Why investors need it:</strong> When Clearbit misses a contact, Hunter's pattern-based inference usually lands it, running both gives near-complete coverage on founder emails at early-stage companies. The deliverability check also prevents the embarrassing bounce on a first intro.</p>
<h2>6. LinkedIn Sales Navigator (paid, from ~$99/mo)</h2>
<p><strong>What it does:</strong> Layers advanced search, saved lead lists, and change alerts on top of LinkedIn. For investors, the killer feature is alerts when a tracked founder changes role, posts, or announces.</p>
<p><strong>Why investors need it:</strong> LinkedIn is still where founders announce fundraises, hires, and pivots. Sales Navigator turns that firehose into a filtered feed of exactly the people you track. Expensive, but for a firm sourcing actively it pays for itself in one caught deal.</p>
<h2>7. BuiltWith (free extension + paid)</h2>
<p><strong>What it does:</strong> Click the icon on any website and it lists the tech stack behind it, hosting, analytics, frameworks, payment providers. The free extension covers most diligence needs; paid plans add historical stack changes.</p>
<p><strong>Why investors need it:</strong> Tech stack is a diligence shortcut. A startup &quot;scaling fast&quot; on a $5/month shared host is not scaling; a team that just migrated off off-the-shelf billing may be hitting real volume. BuiltWith answers &quot;is the engineering real?&quot; in seconds, and pairs naturally with the GitHub lookup in #2 for a full technical-health read.</p>
<h2>8. Save to Notion (free)</h2>
<p><strong>What it does:</strong> One-click clipper that saves any page, a Crunchbase profile, a LinkedIn post, a GitHub repo, into a Notion database with source URL, title, and screenshot.</p>
<p><strong>Why investors need it:</strong> Deal flow is a capture problem before it's an analysis problem. Save to Notion turns browser research into a structured pipeline: every interesting company lands in your tracking database without copy-paste. Pairs with a simple Notion CRM template for anyone not ready to pay for Affinity.</p>
<h2>9. Loom (free tier + paid)</h2>
<p><strong>What it does:</strong> Records screen and camera in one click and produces a shareable video link. The free tier covers 25 videos per account.</p>
<p><strong>Why investors need it:</strong> Partner meetings run on context, and a two-minute Loom walking through a Crunchbase profile and a GitHub repo carries more signal than a one-page memo. Analysts pitch deals upstream with it; partners give feedback asynchronously.</p>
<h2>10. Grammarly (free tier + paid)</h2>
<p><strong>What it does:</strong> Real-time grammar, clarity, and tone checking in every browser text field, Gmail, LinkedIn messages, your CRM.</p>
<p><strong>Why investors need it:</strong> A typo in a cold intro to a technical founder reads as &quot;didn't care enough to proofread.&quot; The free tier eliminates the embarrassing errors; Premium's tone adjustment helps first-touch emails read confident rather than canned. The lowest-glamour item on this list, and the one used most often per day.</p>
<h2>How to combine them: the lean 2026 stack</h2>
<p>You don't need all ten. The leanest high-signal setup for a solo GP or small fund:</p>
<ol><li>Both GitDealFlow extensions, the engineering-momentum signal on Crunchbase, Wellfound, and GitHub.</li><li>Hunter (or Clearbit), for the founder's email.</li><li>Save to Notion, for capturing companies into a pipeline.</li><li>Loom, for pitching deals to partners.</li><li>Grammarly, so the outreach doesn't read like spam.</li></ol>
<p>That stack is free or near-free and covers the full loop: discover → research → capture → reach out → follow up. Firms with budget layer Affinity for relationship tracking and Sales Navigator for LinkedIn signal on top, with BuiltWith for technical diligence on anything approaching a term sheet.</p>
<h2>The signal nobody else has</h2>
<p>The pattern across this list: most VC tools organize or enrich existing data. Affinity organizes your relationships. Clearbit enriches contacts. Sales Navigator filters LinkedIn. All of them work with what's already public and already lagging.</p>
<p>The two GitDealFlow extensions are the only entries that generate a new leading indicator, engineering acceleration drawn from public GitHub activity, benchmarked per-org against its own baseline, and shown in the panel data to precede fundraise announcements by three to six weeks <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>. That's the lead time that lets you start a relationship before the round is competitive.</p>
<p>If you're not ready to install anything, the lowest-friction way to see the signal is the free Sunday digest, one email a week with the startups whose engineering accelerated, at <a href="https://gitdealflow.com">gitdealflow.com</a>. Both extensions install in one click from the <a href="https://signals.gitdealflow.com/install">/install</a> page.</p>
<h2>How to cite this guide</h2>
<p>The Data Nerd (2026). &quot;The 10 Best Chrome Extensions for VC Deal Flow (2026).&quot; VC Deal Flow Signal blog. Retrieved from https://signals.gitdealflow.com/blog/best-chrome-extensions-vc-deal-flow-2026.</p>]]></content:encoded>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[The Startup Due Diligence Checklist: What to Check Before You Write the Check]]></title>
      <link>https://signals.gitdealflow.com/blog/startup-due-diligence-checklist-for-investors</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/startup-due-diligence-checklist-for-investors</guid>
      <description><![CDATA[A working startup due diligence checklist for early-stage investors: team, market, product, financials, and the engineering signals public GitHub data adds before the data room opens.]]></description>
      <content:encoded><![CDATA[<p>Most seed-stage diligence failures do not come from checking the wrong things. They come from checking the right things in the wrong order, and from skipping the engineering layer that is the cheapest, most honest signal available before a data room exists.</p>
<p>This is a working checklist, ordered cheapest-first so the expensive legal work happens only on deals that already survived the public layer.</p>
<h2>Why Order Matters in Due Diligence</h2>
<p>Diligence is a funnel, not a list. Every hour spent on legal review of a company you will pass on is an hour you did not spend sourcing the next deal. Front-loading the free, public checks filters the funnel before you pay for the expensive ones.</p>
<p>The public layer answers three questions no reference call can answer as cleanly: is the team real, is the product real, and is the team shipping. Reference calls tell you what people say; the public record tells you what they do.</p>
<h2>The Four Workstreams</h2>
<p><strong>Team.:</strong> Who are the founders, what have they built before, and does their stated expertise match their track record. Check LinkedIn, GitHub, and prior employers. For technical founders, public commit history is a direct record of capability.</p>
<p><strong>Market.:</strong> Is the market real and growing, or is this a solution looking for a problem. Check comparable companies, funding in the category, and whether the wedge is credible.</p>
<p><strong>Product.:</strong> Does the product exist beyond a deck. Check for a live site, a working demo, real users, and public repositories that corroborate the product story.</p>
<p><strong>Financials and legal.:</strong> Cap table, incorporation documents, prior financing terms, and any outstanding obligations. This is where the NVCA model legal documents come in for priced rounds <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<h2>Where Engineering Data Fits</h2>
<p>Public GitHub activity sits between the team and product workstreams and strengthens both. The GitDealFlow methodology reads commit velocity, contributor growth, and repository expansion across a weekly panel of 350+ startup GitHub orgs <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>. The core finding, backtested against 219 fundraises, is that a composite of commit velocity and contributor growth precedes Series A announcements by 21 to 47 days with a 3.4x lift <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>In diligence terms, that means you can read whether a team is accelerating or stalling before they ever send you a data room link.</p>
<h2>The Five Things Most Investors Skip</h2>
<ol><li>Independent shipping verification, as covered above.</li><li>Contributor breadth. A single founder committing to a single repository is a very different company from a team of eight shipping across five repositories, even at the same raw commit count.</li><li>Repository segmentation. Are the active repositories product code or configuration and policy? Compliance and infra work can inflate activity without advancing the product.</li><li>The founder's own GitHub history. Starring and commit patterns reveal what the founder actually pays attention to.</li><li>Change over time. A snapshot tells you where the company is; the trajectory tells you where it is going.</li></ol>
<h2>A Working Checklist</h2>
<p>Run these in order, and stop as soon as a deal fails two checks in a row. Public presence and team reality first, market second, product third, shipping trajectory fourth, financials and legal last. Each check should end with a written one-line verdict so the decision is reviewable later.</p>
<p>Engineering acceleration is not a substitute for any of the four workstreams. It is the fifth stream that connects team and product, and the one that costs nothing to check before the data room opens.</p>]]></content:encoded>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Technical Due Diligence With Public GitHub Data: Reading Engineering Health Before the Data Room]]></title>
      <link>https://signals.gitdealflow.com/blog/technical-due-diligence-with-github-data</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/technical-due-diligence-with-github-data</guid>
      <description><![CDATA[How to run technical due diligence on a startup using public GitHub activity: commit velocity, contributor growth, repository expansion, and the signals that precede a fundraise.]]></description>
      <content:encoded><![CDATA[<p>By the time a startup sends you the data room, you have already made most of the decision. The data room confirms what you believe; it rarely changes your mind. The technical facts that would change your mind are public earlier, in the GitHub activity the team generates every day.</p>
<p>This is a method for reading that public layer as a pre-data-room technical due diligence pass.</p>
<h2>Why the Data Room Is Too Late</h2>
<p>The data room arrives after the founder has decided to raise, often after the round is already forming. By then the signal you most want, whether the team has been accelerating or stalling, is already priced into the conversation. Public GitHub activity lets you observe the same trajectory in real time, before the round is announced <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>What Public GitHub Reveals</h2>
<p>Three primary signals, each reading a different dimension of engineering health <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>:</p>
<p><strong>Commit velocity:</strong> measures shipping pace across a 14-day window. It is the raw throughput of the team.</p>
<p><strong>Contributor growth:</strong> measures whether the team is expanding or contracting. A jump in active authors is usually capital-driven: the company raised, is deploying capital, and is hiring.</p>
<p><strong>Repository expansion:</strong> measures whether the product surface is broadening. New repositories signal new modules, integrations, or a platform bet.</p>
<h2>Reading the Signals Together</h2>
<p>None of the three means much alone. The composite is what predicts. A team with flat velocity but rising contributor count is onboarding engineers whose output has not landed yet. A team with rising velocity and flat contributors is shipping faster with the same headcount, which is often the strongest short-term signal. A team with rising velocity, rising contributors, and new repositories is executing a coordinated expansion, the pattern that most reliably precedes a priced round <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<h2>The Four Signal Types</h2>
<p>GitDealFlow classifies each startup's activity into a signal type. Engineering hiring bursts indicate a team deploying capital into headcount. Infrastructure buildout indicates a company laying the platform foundation before a growth push. Framework migration indicates modernization. Deploy-frequency spikes indicate a product release cycle. Each points at a different underlying event <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<h2>Where the Method Stops Being Reliable</h2>
<p>Public GitHub data is a health check, not a verdict. It cannot see private repositories, and some of the most important engineering work happens there. It cannot assess code quality, architecture, or security. It can be inflated by compliance and configuration work if you do not segment repositories. And at pre-seed, a thin public footprint means the signal is weak confirmation rather than discovery <sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup>.</p>
<p>Use it to filter and to move faster, not to replace the code review a priced round still deserves.</p>]]></content:encoded>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Enterprise SaaS GitHub Signal Patterns: A Sector Taxonomy for VC Sourcing]]></title>
      <link>https://signals.gitdealflow.com/blog/enterprise-saas-github-signal-patterns</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/enterprise-saas-github-signal-patterns</guid>
      <description><![CDATA[Enterprise SaaS startups signal differently on GitHub than developer-tools or AI companies. A sector taxonomy covering integration API buildouts, SDK releases, contributor compression reversal, and the compliance-cycle false positive, with stage-specific benchmarks from the 350+-startup panel.]]></description>
      <content:encoded><![CDATA[<p>Enterprise SaaS is the sector where the simplest version of the GitHub engineering acceleration model breaks down first. The 14-day commit velocity threshold that works reliably for AI and developer-tools companies systematically under-ranks enterprise SaaS startups, not because the companies are uninteresting, but because the sector carries a structurally lower commit baseline that a naive percentile ranking treats as inactivity.</p>
<p>Across the 350+-startup panel <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>, enterprise SaaS companies average 35-60% fewer commits per 14-day window than AI-tools startups at equivalent funding stage. The causal chain is short: SOC 2 Type II compliance requirements, enterprise customer change-notification windows, and the review gates embedded in most B2B engineering processes all throttle deployment frequency in ways that public GitHub commit data reflects directly. DORA research confirms that deployment frequency is significantly lower in regulated and customer-sensitive environments than in product-led growth companies <sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup>. A two-week sprint that produces 20 commits in a developer-tools company may produce 12 commits at an enterprise SaaS company doing equivalent engineering work.</p>
<p>The implication for signal extraction is that threshold-based screens configured for the general startup population will systematically filter out enterprise SaaS breakouts. A company accelerating from 18 commits per 14-day window to 36 commits, a +100% acceleration, looks unimpressive in absolute terms but is, within the enterprise SaaS sector distribution, a top-decile move. Sector-stratified ranking, computing percentile rank within enterprise SaaS rather than across the full panel, is the correct default for this sector <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<h2>Why Velocity is the Wrong Primary Metric</h2>
<p>Commit velocity is the dominant metric in the general engineering acceleration model because it correlates with shipping pace, which correlates with team expansion, which correlates with capital events <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>. This chain holds across most sectors.</p>
<p>In enterprise SaaS, the correlation weakens at the first link. Shipping pace is bounded by customer contracts, not solely by team capability. Adding three engineers to an enterprise SaaS team does not produce a proportional commit increase, it produces a proportional increase in capacity to serve customer integrations, maintain compliance infrastructure, and build features requested by named accounts. That work is substantial and capital-intensive, but a meaningful fraction may be conducted in private repositories, customer-specific branches, or infrastructure work that generates few public commits.</p>
<p>The stronger primary metrics for enterprise SaaS are contributor count change and repository expansion. These are less bounded by compliance constraints than velocity: hiring is observable in contributor graphs regardless of deployment pace, and repository creation reflects strategic decisions that are independent of sprint-level noise. The practical model for this sector is to treat commit velocity as a secondary confirmation metric and treat contributor count and repository count as the primary leading indicators.</p>
<h2>The Five Dominant Signal Patterns</h2>
<p>Five patterns account for the majority of actionable enterprise SaaS GitHub signals across the panel.</p>
<p><strong>Integration API buildout:</strong> is the most reliable fundraise precursor in this sector. The pattern: a cluster of new repositories appears within a 20-to-35-day window, each named after a major enterprise platform, Salesforce connectors, HubSpot integrations, Workday adapters, Slack applications. The cluster indicates that the company has enough enterprise customers requesting specific integrations to justify parallel product investment. By the time a company is building four integrations simultaneously, the product-market loop in the enterprise channel is established, the fundraise to scale it follows 4-7 weeks later in the panel data <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>. Detection rule: at least three new integration-named repositories created within a 30-day window, each showing sustained commit activity above 15 commits in the first 14 days after creation.</p>
<p><strong>SDK and CLI release:</strong> is the second-strongest pattern. A public SDK release signals that the company is confident its API is stable enough to invite external developers to depend on it, a platform bet that typically follows a strong Series A or substantial customer traction. SDK repositories are detectable by naming convention, sdk-, -client, -cli, -py, -js suffixes, and by a distinctive contributor pattern: internal contributors dominate the first four weeks, then external stars and forks accumulate as the market responds <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>. The fundraise signal is strongest when the SDK release is accompanied by a new CNAME or subdomain detectable via certificate transparency logs within the same window.</p>
<p><strong>Multi-module expansion:</strong> is the architectural signal. Rather than a single new repository, multi-module expansion produces five to ten new repositories within 30-45 days, each representing a separable product component: an admin API, a data-export service, a webhook delivery layer, a customer portal, an embeddable widget. This decomposition almost always reflects enterprise customer requirements for customization and independent deployment. It is most diagnostic when the new repositories span different language stacks, TypeScript on the frontend, Go or Rust on the backend, indicating deliberate system design rather than placeholder scaffolding. Multi-module expansion at the Series A-to-B inflection has the highest correlation with follow-on funding of any pattern in the enterprise SaaS subset of the panel.</p>
<p><strong>Contributor compression reversal:</strong> exploits the sector's tendency to maintain small, stable engineering cores. Many enterprise SaaS startups operate with two to five active contributors for 9-15 months, building a core product to the point of first enterprise customers. When contributor count jumps from that compressed state, from four to nine contributors in a single 28-day window, for example, the move is almost always capital-driven rather than community-driven. This makes contributor compression reversal a cleaner hiring-burst signal than in developer-tools or AI, where external community contributors can produce similar jumps unrelated to the company's financials. Detection rule: a 75% or greater increase in unique 28-day contributing authors from a prior stable baseline of five or fewer.</p>
<p><strong>The compliance-to-product pivot:</strong> is the subtlest pattern. Some enterprise SaaS startups begin their GitHub life as primarily compliance-and-configuration repositories, infrastructure as code, audit logs, security tooling, and then shift toward product repository creation as the core platform matures. This pivot is observable as a change in repository type mix: the ratio of new product-facing repositories to new infrastructure repositories increases sharply over a quarter. A startup with eight infrastructure repos and two product repos that adds four product repos and zero infrastructure repos in a 45-day window has pivoted its engineering investment toward customer-facing features, a reliable pre-fundraise signal regardless of absolute velocity.</p>
<h2>Stage-Specific Benchmarks</h2>
<p>At **pre-seed**, enterprise SaaS GitHub signals are weak discovery tools but reasonable verification tools. Typical pre-seed organizations have one to four public repositories, two to five contributors, and 10-35 commits per 14-day window. A team at the high end of this range is meaningfully differentiated from the median but still insufficient for confident cold discovery. Use pre-seed enterprise SaaS signals to confirm a team encountered through another channel, not to discover one independently.</p>
<p>At **seed**, the integration API buildout pattern first becomes diagnostic. A seed-stage team with one to three integration repositories and five to ten contributors showing contributor compression reversal, moving from a stable prior baseline to a 75%+ increase, should be treated as a high-priority signal. The typical seed-stage breakout shows this transition within 6-10 weeks of a round close, making it an investable lead rather than a retrospective observation.</p>
<p>At **Series A**, multi-module expansion is the dominant pattern. A team accelerating from five repositories to twelve to eighteen over a single quarter, with sustained commit activity across the new repos, is building platform infrastructure in response to enterprise customer requirements. Repository expansion at this rate is uncommon at Series A in any other sector, which gives it relatively high specificity as an enterprise SaaS indicator.</p>
<p>At **Series B and later**, single-metric breakouts are almost always explainable by organizational events, an acquisition, a team merger, a monorepo migration, and require cross-validation before investment action. The useful signal at this stage is composite: velocity, contributor count, and repository count moving together over a 28-day window, with no obvious structural explanation from public company announcements.</p>
<h2>False Positives: The Enterprise SaaS Checklist</h2>
<p>Three false-positive categories are disproportionately common in enterprise SaaS.</p>
<p><strong>Compliance-cycle bursts.:</strong> SOC 2, ISO 27001, and FedRAMP renewal cycles produce predictable engineering bursts in audit, policy, and infrastructure repositories every 12 months. The diagnostic check is repository segmentation: identify which repositories are accelerating. Acceleration concentrated in repositories named or structured as compliance tooling should be discounted. Concurrent acceleration in product or integration repositories, even if smaller in absolute magnitude, changes the interpretation and may be worth pursuing.</p>
<p><strong>Customer-specific branch activity.:</strong> Some enterprise SaaS teams maintain customer-specific forks or branches in public repositories, which inflate commit counts around onboarding events. The diagnostic check is author email domain concentration: inspect commit metadata via the GitHub REST API <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>. A cluster of commits originating from a single external company domain indicates customer onboarding work rather than internal product investment.</p>
<p><strong>Legacy migration noise.:</strong> Platform migrations, from one cloud provider to another, from monolith to microservices, produce explosive repository creation and velocity that reflects architectural transition rather than growth. The diagnostic check is whether new repositories appear to supersede existing ones: names including -v2, -new, or -next adjacent to a simultaneous commit drop in older repositories indicate migration rather than expansion.</p>
<h2>Cross-Validation Layers</h2>
<p>Enterprise SaaS GitHub signals are meaningfully stronger when combined with two additional data sources. Job postings for enterprise sales roles, Account Executive, Solutions Engineer, Enterprise Customer Success, appearing concurrently with a contributor compression reversal are nearly diagnostic: the combination indicates a company simultaneously scaling engineering and go-to-market, the canonical Series A-to-B transition pattern. Second, domain-level infrastructure changes, a new api.company.com or enterprise.company.com subdomain detectable via certificate transparency logs, within the same window as a multi-module GitHub expansion confirms that the product surface has reached the customer-facing layer.</p>
<p>Neither cross-validation source eliminates the need for a founder conversation, but together with the GitHub signal, they compress the false-positive rate for enterprise SaaS to a workable level, consistent with the broader panel performance documented in <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>.</p>
<p>The sector-stratified rankings, signal thresholds, and underlying methodology are documented at <a href="https://signals.gitdealflow.com/methodology">signals.gitdealflow.com/methodology</a>. The longitudinal panel underpinning these observations is described in the SSRN working paper at <a href="https://ssrn.com/abstract=6606558">ssrn.com/abstract=6606558</a>.</p>]]></content:encoded>
      <pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[30 Research Findings, Now One Page Each: How to Cite GitHub Engineering Acceleration]]></title>
      <link>https://signals.gitdealflow.com/blog/30-research-findings-now-one-page-each</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/30-research-findings-now-one-page-each</guid>
      <description><![CDATA[Every atomic finding from the SSRN-indexed GitDealFlow paper now lives on its own page with ScholarlyArticle schema, citation chain, and how-to-cite block. Easier to quote, easier to link, easier for AI engines to attribute correctly.]]></description>
      <content:encoded><![CDATA[<p>_VC Deal Flow Signal (GitDealFlow). On this site, &quot;engineering acceleration&quot; means a quantitative GitHub momentum signal, not a reference to startup accelerator programs._</p>
<h2>What changed</h2>
<p>The 30 atomic findings from our SSRN-indexed paper used to live on a single long page at <a href="https://signals.gitdealflow.com/research">/research</a>. Today every finding has its own URL, its own ScholarlyArticle JSON-LD entry, its own citation block, and its own slug.</p>
<p>Cite the median commit velocity directly:
<a href="https://signals.gitdealflow.com/research/median-commit-velocity-venture-startups">/research/median-commit-velocity-venture-startups</a></p>
<p>Cite the 75% framework-migration share directly:
<a href="https://signals.gitdealflow.com/research/framework-migration-dominant-signal-type">/research/framework-migration-dominant-signal-type</a></p>
<h2>Why split</h2>
<p>A single long page is a bad citation target. Every quotable number now resolves to its own URL, which means AI engines and citation tools can deep-link with confidence and the answer-engine attribution model finally has somewhere to point.</p>
<h2>What each sub-page carries</h2>
<p>- **ScholarlyArticle JSON-LD** with the headline, abstract (the &quot;why it matters&quot; line), citation→SSRN, sameAs chain to OpenAlex/Crossref/Zenodo
- **BreadcrumbList** for SERP breadcrumb display
- **Speakable** selector on H1 for voice-assistant extraction
- **Provenance block** linking to the SSRN paper, the dataset DOI, the CC BY 4.0 license, the author ORCID, and the Wikidata Q-item
- **How-to-cite block** with copy-paste citation in plain-text form
- **Prev/next navigation** so search-arrived users can browse adjacent findings</p>
<h2>Internal-link wiring</h2>
<p>The /research index page now links each finding to its sub-page. The homepage carries a six-finding cluster transferring PageRank from the highest-DA page on the site. The sitemap (split-index format) carries 31 new URLs in <a href="https://signals.gitdealflow.com/sitemap/content.xml">/sitemap/content.xml</a>. The qa.jsonl corpus at <a href="https://signals.gitdealflow.com/qa.jsonl">/qa.jsonl</a> gained 30 new Q&amp;A entries, one per finding.</p>
<h2>How to find the rest</h2>
<p><a href="https://signals.gitdealflow.com/research">/research</a> is the index. Every finding card on that page is a link to its sub-page. The full cross-graph identity map (every external anchor, Wikidata, ORCID, SSRN, OpenAlex, Crossref, Semantic Scholar, Zenodo, DataCite, code repositories, social profiles) is at <a href="https://signals.gitdealflow.com/citations">/citations</a>.</p>
<h2>How to cite this announcement</h2>
<p>The Data Nerd (2026). &quot;30 Research Findings, Now One Page Each.&quot; VC Deal Flow Signal blog. Retrieved from https://signals.gitdealflow.com/blog/30-research-findings-now-one-page-each.</p>
<p>Replication studies welcome. signals@gitdealflow.com for co-authorship on funding-event joins.</p>]]></content:encoded>
      <pubDate>Fri, 01 May 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[How VCs Track Startup Engineering Acceleration: The Complete 2026 Playbook]]></title>
      <link>https://signals.gitdealflow.com/blog/how-vcs-track-engineering-acceleration-2026-playbook</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/how-vcs-track-engineering-acceleration-2026-playbook</guid>
      <description><![CDATA[The complete 2026 playbook on engineering acceleration as a VC deal flow signal, pipeline, metrics, benchmarks, predictive analytics, screening workflow, and sector patterns, with worked examples from a 350+-startup GitHub panel.]]></description>
      <content:encoded><![CDATA[<p>Engineering acceleration, the rate of change in a startup's GitHub commit, contributor, and repository activity relative to its own 14-day baseline, is one of the highest-signal leading indicators investors can read today. Across the 350+-startup panel maintained at VC Deal Flow Signal <sup><a href="https://signals.gitdealflow.com/blog/#references">[4]</a></sup><sup><a href="https://signals.gitdealflow.com/blog/#references">[5]</a></sup>, a sustained doubling in 14-day commit velocity precedes a public fundraise announcement by a median of three to six weeks.</p>
<p>That window matters. By the time a startup appears in Crunchbase, PitchBook, or even press coverage, the round is often allocated and competitive. Engineering acceleration shows up before the funding databases catch up, before the press cycle, and often before the founder has even sent the first investor email.</p>
<p>This playbook is the operating model behind that observation. It covers the full pipeline: how to collect the GitHub data, which metrics to derive, how to benchmark them, how to turn the signal into a fundraise forecast, how to automate screening at fund scale, how to integrate the output into an existing sourcing workflow, and where the model breaks down. The framework is deliberately reproducible, the underlying data is public, the metrics are simple, and any analyst comfortable with REST APIs can build a working version in a week.</p>
<p>A vocabulary note before going further. Throughout this piece, engineering acceleration refers exclusively to code-side momentum: a measurable change in shipping pace observable in public commits, contributors, and repositories. It has nothing to do with startup accelerator programs like Y Combinator or Techstars, despite the unfortunate vocabulary collision.</p>
<h2>Why engineering acceleration is the leading indicator most VCs missed</h2>
<p>For decades the venture capital industry has converged on a small set of sourcing channels: founder networks, demo days, accelerator pipelines, and increasingly the proprietary databases sold by PitchBook, Crunchbase, Affinity, and a handful of newer entrants. Each of those surfaces has the same fundamental limitation: it is downstream of the actual product. The data appears after a press release, after a board introduction, after a CRM update from a partner.</p>
<p>Engineering acceleration is upstream of all of those signals. Code is the earliest publicly observable artifact of a working startup. Before there is a website redesign, before there is a hiring announcement, before there is a Crunchbase entry, there are commits to a public repository. For technical startups, meaning any company whose core product or platform is built on software, which includes the majority of venture-backed companies in 2026, the GitHub footprint is the leading edge of all other observable activity.</p>
<p>The reason this surface has been historically under-monitored is not that the data is hidden. GitHub's REST and GraphQL APIs <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup> expose commit activity, contributor lists, repository metadata, language statistics, and release history. A solo developer can pull a year of commit history for any public organization in a few seconds. The constraint has been engineering effort: building a longitudinal panel across thousands of startups, normalizing for noise, computing rate-of-change metrics, and surfacing breakouts requires a small but specialized data engineering team. Until recently, no one in the venture industry had built it at scale.</p>
<p>The economic case is straightforward. A fund that can identify a Series A-bound startup six weeks before it appears in Crunchbase Alerts has a structural sourcing advantage over peer funds limited to the same downstream sources. The advantage compounds: warm outreach during the pre-fundraise window converts at meaningfully higher rates than cold outreach during a competitive round. Even at the angel and seed level, where check sizes are smaller and the marginal return on each deal is bounded, the ability to be the first thoughtful conversation a founder has had about a round is worth disproportionately more than being the tenth.</p>
<p>The signal is not magic. Many accelerations resolve in disappointing ways: the team ships a launch and stalls, the round is extended rather than upsized, the apparent engineering burst was a single contributor's hackathon. The framework's job is not to eliminate those false positives, it is to make them tractable, to reduce the screening surface area from every venture-backed startup to the 30 to 50 companies in your sectors of interest currently showing breakout engineering activity. That is a workload a single analyst can clear in a Monday morning.</p>
<p>The remainder of this playbook is the operating manual.</p>
<h2>Building a GitHub data pipeline for VC deal flow</h2>
<p>The first practical question is how to acquire and organize the data. The pipeline has four layers: target list, ingestion, normalization, and aggregation.</p>
<p>The target list is a curated set of startup GitHub organizations. There is no canonical source. The pragmatic approach is to seed the list from existing investor lists, Y Combinator's published company directory, AngelList's startup database, sector-specific accelerator alumni, and union them with discoveries surfaced through the pipeline itself. At VC Deal Flow Signal, the working list is approximately 350+ organizations covering 15 sectors. The exact size is less important than the maintenance discipline: the list must be reviewed monthly, with stale entries pruned and new entries added based on press, demo days, and reader submissions.</p>
<p>Ingestion uses the GitHub REST API <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup>. The free authenticated rate limit is 5,000 requests per hour per token. For a 350+-organization panel, the relevant endpoints are repositories list, commit activity, contributors, and releases. A weekly full pass requires roughly 50,000 requests, which is comfortably within the rate limit if the work is split across two tokens or paced over six hours. The GitHub Innovation Graph dataset <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup> provides aggregate cross-region statistics that are useful for sector-level benchmarks but not for individual startup tracking.</p>
<p>Normalization is the most overlooked stage. Raw commit counts are noisy: bot accounts, dependency-update PRs from Dependabot or Renovate, automated formatting commits, and CI/CD reruns can inflate counts without reflecting real engineering activity. The minimum viable normalization is excluding commits authored by accounts whose name matches common bot patterns. A more sophisticated pipeline classifies commits by file diversity, message length, and diff size. Each layer of normalization improves signal-to-noise but adds engineering cost; the practical sweet spot is bot exclusion plus file-count filtering, which removes the loudest noise sources without overfitting.</p>
<p>Aggregation produces the metrics that drive the rest of the playbook. The minimal set is four time series per organization: commit velocity (count over rolling 14 days), unique contributor count over the same window, repository expansion (new repos created in the period), and language mix as a percentage breakdown of commit volume by primary language. Each is computed weekly and stored with timestamps so that subsequent comparisons against historical baselines are exact.</p>
<p>A crucial design choice is the comparison window. Investor signal pipelines tend to use either 14-day rolling or 28-day rolling windows. The 14-day window is more responsive, it surfaces breakouts faster, at the cost of higher volatility. The 28-day window is smoother but introduces lag. The pragmatic default is 14-day with a confirmation rule: a breakout must persist into a second 14-day window before it is treated as actionable. This filter alone removes most one-period spikes caused by hackathons, launch sprints, or single contributors onboarding.</p>
<p>The pipeline output is a per-organization weekly snapshot: four metrics, four prior-period baselines, and four rate-of-change values. Stored in a flat database, the entire 350+-organization panel fits in well under a gigabyte. The data infrastructure does not require dedicated streaming infrastructure or a data warehouse, a single Postgres or DuckDB instance handles the volume comfortably.</p>
<p>The pipeline is, in short, an unglamorous engineering project: rate-limited API calls, careful joins, and a weekly cron. The competitive moat is not in any one component but in the months of accumulated panel history that newer entrants cannot replicate retroactively.</p>
<h2>Core metrics: commit velocity, contributor growth, and repository expansion</h2>
<p>Engineering acceleration as a single number is useful for ranking, but the underlying signal decomposes into four metrics, each of which carries a different operational meaning.</p>
<p>Commit velocity is the headline metric: total commits to the most active repositories of a startup organization over a 14-day window. Velocity correlates with team size, codebase maturity, and shipping cadence. Absolute velocity is not directly comparable across companies, a 10-person team will out-commit a solo founder by a factor of 10 even at equivalent per-engineer pace, but velocity change relative to a startup's own baseline is the most honest comparable. A startup that has gone from 30 commits per period to 70 has demonstrably accelerated; whether that move was driven by hires, a sprint, or simply higher per-engineer output is downstream interpretation.</p>
<p>Contributor growth is the second-order signal. Velocity can rise because existing engineers are working harder; contributor count can only rise when new engineers are added. Tracking unique authoring contributors over a 14-day window, and comparing to the prior window, distinguishes the existing team is sprinting from the team has hired. The latter is a stronger fundraise signal because hiring almost always implies committed capital, while sprinting can occur within an existing runway. A typical pre-Series-A pattern looks like contributor count moving from three to six to nine over consecutive 14-day windows, often within four to eight weeks of the round close.</p>
<p>Repository expansion captures structural changes. New repositories appearing in a startup's organization signal product line extension, infrastructure migration, or in some cases a pivot. The pattern is most informative in conjunction with velocity: a new repo with sustained commits over its first month is a launch precursor; a new repo with a single commit that goes silent is usually a placeholder. The signal works best for organizations with a stable repository naming convention, where the addition of a repo named after a new product surface is recognizable without context.</p>
<p>Language mix evolution is the slowest-moving but most strategic signal. A team whose commit volume shifts measurably from Python to TypeScript, or from JavaScript to Rust, has changed something fundamental about the system being built. These transitions are infrequent, typical startups maintain stable language mixes over months, but when they occur they often coincide with platform rebuilds, performance milestones, or hires of senior engineers with specific stack preferences. A 20-percentage-point shift in language mix over a single quarter is unusual and merits investigation.</p>
<p>The four metrics combine into a typology of acceleration patterns. The hiring burst is contributor count plus velocity rising together. This is the pattern most strongly correlated with a recent or imminent fundraise. Detection rule: contributor count up at least 30 percent and commit velocity up at least 60 percent in the same period.</p>
<p>The shipping sprint is velocity rising while contributor count holds flat. This signals a launch push, often around a specific product milestone. Detection rule: velocity up at least 100 percent with contributor count change under 15 percent.</p>
<p>The infrastructure buildout is repository creation accelerating relative to historical baseline. This signals architectural investment, often platform migrations or a build-out of internal tooling. Detection rule: at least three new repositories created in 30 days versus a prior 30-day baseline of zero.</p>
<p>The platform migration is language mix shifting. This is the slowest-moving pattern but often the most strategically significant, it implies the team is committing to a new technical direction. Detection rule: at least 20 percentage points of language mix migrating between primary languages over a quarter.</p>
<p>Each pattern has implications for diligence. A hiring burst suggests asking about recent or imminent capital. A shipping sprint suggests asking about the upcoming launch and its dependencies. An infrastructure buildout suggests asking about architectural strategy and the team's view of scaling. A platform migration suggests asking about the technical bet driving the change. The metrics direct the investor's attention; the actual decision still requires founder conversations.</p>
<h2>Benchmarking: what healthy acceleration looks like at each stage</h2>
<p>Raw acceleration numbers without context are meaningless. A pre-seed team with two contributors and a +200% commit velocity change is showing different physics than a Series B company with 50 engineers and a +200% change. Benchmarking against stage and sector is essential to interpret signals correctly.</p>
<p>The pre-seed benchmark is dominated by base rate effects. A solo founder going from 5 commits in a 14-day window to 20 shows +300%, but the absolute volume is too low to draw structural conclusions. The useful pre-seed signal is sustained activity: contributor count holding above 1, velocity holding above 10 commits per 14-day window for at least eight weeks, and at least one external dependency or release tag indicating production-grade engineering. At pre-seed, the goal of the metric is to identify teams that are credibly building, not to compare them.</p>
<p>The seed benchmark introduces structural comparisons. A typical seed-stage team is 3 to 8 engineers committing 80 to 200 commits per 14-day window. A breakout signal at seed is a sustained move above the 75th percentile of the company's prior six months, typically a velocity increase of 50 percent or more held for at least 28 days. Contributor count change at this stage is highly informative: a seed team going from 5 to 8 engineers in a quarter is making a real bet that almost always reflects either a fundraise or imminent fundraise.</p>
<p>The Series A benchmark shifts toward magnitude. A Series A team commits 200 to 500 commits per 14-day window across a larger codebase. The interesting signal at this stage is not whether a team is accelerating in percentage terms, many teams maintain steady acceleration as they hire, but whether the acceleration involves new product surfaces. A Series A team that adds a new repository with sustained activity, or that visibly migrates to a new primary language, is signaling strategic shifts that often precede follow-on fundraises or major launches.</p>
<p>The Series B and later benchmark is dominated by composition. Acceleration at this stage is often driven by acquisitions, new business unit launches, or the absorption of a technical hire team rather than a singular team push. The signal mix shifts: contributor count is less informative because hiring is structurally embedded; repository expansion becomes more meaningful as it reflects strategic bets; language mix changes can presage M&amp;A activity when the absorbed team's stack appears in the parent organization's commit graph.</p>
<p>Sector-specific benchmarks layer on top of stage benchmarks. AI and machine learning startups typically have higher commit volatility than enterprise SaaS startups because model training cycles, dataset releases, and notebook-driven research all produce large bursts of activity. Fintech infrastructure startups tend to have lower commit volume but higher per-commit substance because compliance and security review act as natural throttles. Developer tools startups have the highest contributor diversity because open source contribution is itself part of the product strategy. The full sector breakdown is detailed in the sector patterns section below.</p>
<p>The benchmarking takeaway is that no single threshold works across the population. The +100% rule is a reasonable default for screening, but the ranking that matters is each startup against its own historical baseline plus a sector-stage adjusted control. A working pipeline computes both: an absolute acceleration number for a fast first-pass filter, and a percentile-rank-within-sector-stage for prioritization. Investors looking at the screen can run on either depending on whether they are sourcing for breadth or for conviction.</p>
<h2>Predictive analytics: turning GitHub signals into fundraise forecasts</h2>
<p>A useful sourcing pipeline produces not just rankings but probabilities. For each startup showing acceleration, what is the probability it will announce a fundraise in the next 8 weeks? Translating signals into forecasts requires labeling and calibration.</p>
<p>The labeling problem is harder than it sounds. A fundraise announcement can be a press release, an SEC Form D filing, a Crunchbase entry, a founder tweet, or in some cases simply a quiet update to a company website. Building a reliable label set requires a multi-source ground truth: tagging each company in the panel as having raised within an N-week window using whichever source confirms first. At VC Deal Flow Signal, the labeling pipeline merges Form D filings, Crunchbase, Affinity, and a manual press review for ambiguous cases.</p>
<p>Once labels exist, the predictive model can be straightforward. A logistic regression on the four core metrics, commit velocity change, contributor count change, new repository count, and language mix shift magnitude, over the prior 28 days predicts a near-term fundraise with measurable lift over base rate. More sophisticated models add interactions (velocity-times-contributor-change captures hiring bursts cleanly), sector indicators, and stage indicators. Tree-based models like gradient-boosted trees provide a few percentage points of additional precision but at meaningful interpretability cost. The pragmatic recommendation is to start with logistic regression for transparency and only add complexity once the simple model's failure modes are well understood.</p>
<p>Calibration is critical for usefulness. A model that says 70 percent probability of fundraise needs that number to actually mean what it says when aggregated across many predictions. Reliability diagrams, bucketing predicted probabilities and checking the empirical fundraise rate within each bucket, should be a routine output of any working pipeline. Most investor-facing rankings do not need probabilities at all; they need a defensible relative ordering. But for funds running portfolio-level analyses, calibration matters.</p>
<p>The model's outputs interact with portfolio construction in a specific way. Top-ranked breakouts are not necessarily the most investable signals. A Series B team showing acceleration is usually less actionable for an early-stage fund than a seed-stage team showing similar magnitude, even though the Series B signal is statistically stronger. The screening pipeline should report both raw probability and stage-adjusted rank, allowing each fund to filter to the population that matches their mandate.</p>
<p>There is a meta-question worth flagging. If engineering acceleration becomes a widely used investor signal, will markets adapt? The answer is partly yes, partly no. Founders intentionally trying to game the signal would need to sustain real engineering output across multiple contributors, which is itself a form of real activity. More plausibly, founders aware of the signal will time their public-repository activity to optimize visibility, pushing changes from local branches in coordinated bursts, for example. But the underlying causal structure is difficult to fake without doing the underlying work, and the signal's persistence as a leading indicator depends on this asymmetry.</p>
<p>The probabilistic output, once trusted, has uses beyond ranking. Portfolio funds can monitor existing investments for engineering acceleration, a portfolio company quietly accelerating may be ready for a follow-on conversation. Limited partners in venture funds can use aggregate sector-level acceleration as a leading indicator of where the next vintage of breakouts is forming. Strategic acquirers can use the same data to identify acquisition targets before banker auctions begin.</p>
<p>The forecasting layer is where the engineering effort pays off. The pipeline that produces unranked rankings is useful; the pipeline that produces calibrated probabilities is differentiated.</p>
<h2>Automating screening: workflow at the fund level</h2>
<p>Most investor processes do not benefit from raw data; they benefit from a curated weekly digest. Automating the screening layer is what turns the underlying pipeline into a usable product.</p>
<p>The screening workflow has four stages: signal detection, deduplication, contextualization, and prioritization. Signal detection is the threshold pass: identify all startups exceeding the +100% sustained acceleration threshold or the percentile-rank threshold within their sector. Deduplication removes companies already actively in the fund's pipeline based on a CRM cross-reference. Contextualization attaches enrichment, recent press, hiring activity, founder Twitter, to each surviving signal. Prioritization ranks the result list by a combination of signal strength, sector fit with the fund's mandate, and stage fit.</p>
<p>The output is a Monday-morning report: typically 20 to 40 startups for a fund with focused sector mandates, longer for sector-agnostic generalist funds. Each entry includes the signal type (hiring burst, shipping sprint, infrastructure buildout, platform migration), the signal magnitude relative to baseline, the GitHub link, the team's public profile, and any enrichment data. The report is written for a 15-minute Monday-morning review followed by partner-level triage.</p>
<p>Automating CRM integration is where most funds first encounter friction. Affinity, HubSpot, and Salesforce all expose APIs for cross-referencing organization names against existing pipeline records. The match logic must be tolerant: founders use varied legal names, GitHub organization names, and public-facing brand names. A working CRM cross-reference uses domain-based matching as a primary key and falls back to fuzzy name matching with a manual review queue for ambiguous matches. Building this layer once saves hours of analyst time per week.</p>
<p>Contextualization automates the five minutes per company research the analyst would otherwise do manually. The standard enrichment stack pulls recent press from Google News, hiring activity from LinkedIn or Wellfound, founder social activity from Twitter and Hacker News, and any prior fundraise history from Crunchbase. Each signal type benefits from different enrichments, a hiring burst is most informatively contextualized with LinkedIn headcount data, while a shipping sprint is best contextualized with the team's public roadmap or product page.</p>
<p>Prioritization is where fund-specific judgment lives. A pre-seed-focused angel weights stage fit highly and sector fit lightly. A specialized Series A fund weights both heavily. A platform fund weights signal strength and recency over fit. The screening pipeline must be configurable per fund without requiring every fund to maintain its own infrastructure. The product implication is that the right level of abstraction is ranked weekly digest rather than raw data feed, a fund that wants raw data feeds typically has the engineering capacity to build its own pipeline, and the product is more useful as the weekly summary.</p>
<p>A subtle but important workflow detail is the feedback loop. Funds that mark which signals they actually pursued, and which converted into meetings or investments, provide data that improves the ranking model. A fund-private feedback loop, without sharing data across funds, improves the ranking weights for that fund's specific mandate. This level of personalization is not necessary for the first version of the pipeline but materially improves precision over months of use.</p>
<p>The full automated screening loop runs in a few hours per week of compute. The human-in-the-loop time is roughly 30 minutes Monday morning per analyst. The per-deal cost of incremental sourcing, compared to traditional channels, is competitive even at the smallest fund sizes. The economic argument for automation is not labor savings; it is coverage. A single analyst monitoring engineering acceleration across 350+ organizations would otherwise require an unworkable amount of attention. Automation makes the coverage tractable, and tractable coverage is the actual sourcing edge.</p>
<h2>Workflow integration: how investors operationalize the signal</h2>
<p>The signal is only useful if it lands in the actual investment process. Integration with existing workflow tools is where many promising data products fail in deployment. Three integration patterns work in practice.</p>
<p>The weekly digest pattern is the simplest. The screening pipeline emails a Monday-morning report directly to partners and analysts, formatted for skimming. Each entry has a one-line signal summary, a magnitude, and a one-click link to a deeper view. This pattern requires no infrastructure on the fund side and works well for small funds and angels. Its limitation is that the digest exists outside the fund's CRM and pipeline workflows; tracking which leads converted requires manual logging.</p>
<p>The CRM enrichment pattern extends the screening output into the fund's existing CRM. New signals appear as tagged opportunities in Affinity or HubSpot with the signal type and magnitude as fields. Existing pipeline companies receive enrichment events when their engineering acceleration changes meaningfully. This pattern requires API integration but produces tightly closed loops, every signal lands in the same surface where decisions are tracked, and conversion data flows naturally back to the screening model.</p>
<p>The Slack or email alert pattern handles real-time intervention. Rather than waiting for a weekly digest, alerts fire when a portfolio company crosses a threshold (engineering acceleration above +200%, contributor count growing faster than expected, infrastructure buildout signals in a competitor's organization). This pattern is most useful for funds with active portfolio management practices or for competitive intelligence against other funds' portfolios.</p>
<p>The most operationally mature funds combine all three patterns: weekly digest for top-of-funnel surfacing, CRM enrichment for closed-loop tracking, and real-time alerts for high-priority events. The complexity escalates accordingly, and the effort is not justified for the first months of use. The pragmatic deployment path is to start with weekly digest for two months, add CRM enrichment once signals are converting reliably, and add real-time alerts only after the first two layers have demonstrated value.</p>
<p>A frequently raised concern is signal fatigue. Funds receiving 30 to 40 signals per week can develop dismissal habits if the signal-to-noise ratio is weak. Two countermeasures help. First, ranking discipline: never present more signals than the analyst can review thoughtfully in 30 minutes. Second, conversion tracking: explicitly track how many signals converted to first meetings, meetings to diligence, and diligence to investment. Funds that see 5 percent of signals convert to first meetings, and 1 percent eventually convert to investments, are running a workable funnel. Funds that see no conversion in three months should investigate either the screening pipeline's calibration or the fund's sector fit.</p>
<p>A subtlety that deserves explicit treatment is the relationship between this signal and the founder. The most successful fund deployments treat engineering acceleration as a conversation prompt, not a buying signal. The first outreach is, by design, low-pressure: the fund has noticed that the team is shipping unusually fast, would like to learn what is driving it, and is open to whatever the founder wants to share. Founders almost universally respond positively to this framing because it acknowledges their work and does not presume a transaction. Funds that lead with a transaction-focused outreach often see lower response rates because the founder has not yet decided to raise.</p>
<p>Workflow integration is, ultimately, a discipline question. The fund that treats the signal as a structured input to a normal investment process gets compounding value. The fund that treats it as a one-off curiosity gets a few interesting Monday morning emails and not much more.</p>
<h2>Sector patterns: how acceleration looks in AI, fintech, devtools, and other technical verticals</h2>
<p>Engineering acceleration manifests differently across sectors. Understanding sector-specific patterns is the difference between a noisy first-pass screen and an interpretable signal.</p>
<p><a href="https://signals.gitdealflow.com/blog/how-vcs-track-engineering-acceleration-2026-playbook">Read the full analysis</a></p>]]></content:encoded>
      <pubDate>Sun, 26 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[I made my VC deal flow callable by Claude this weekend. Here is what that actually means.]]></title>
      <link>https://signals.gitdealflow.com/blog/a2a-launched</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/a2a-launched</guid>
      <description><![CDATA[GitDealFlow now publishes a Google A2A AgentCard at /.well-known/agent-card.json and a JSON-RPC 2.0 endpoint at /api/a2a. Five free skills, no auth. Crunchbase API costs $20K per year. Ours costs nothing. Here is the curl that proves it.]]></description>
      <content:encoded><![CDATA[<p>Saturday afternoon. I was in Cursor running a thread about a fintech startup my Telegram channel had flagged that morning. Cursor was good at the GitHub code review, fine at reading the README, helpful with the package.json. Then I asked it: &quot;How does this company's commit velocity compare to other fintech startups this quarter?&quot;</p>
<p>It could not answer. Not because the data did not exist. We publish it weekly at <a href="https://signals.gitdealflow.com">signals.gitdealflow.com</a>. It could not answer because there was no way for it to ASK us programmatically the way I would ask a colleague.</p>
<p>My AI had hands. It just did not have a phonebook.</p>
<h2>The $20,000 phonebook</h2>
<p>When you peel back what every AI-powered VC tool actually does, it is the same loop: a human types a question, an LLM tries to answer, the LLM hits a wall because the data is not there, the LLM apologizes. Every. Time.</p>
<p>The fix is APIs. The problem is APIs:</p>
<p>- Crunchbase API: $20,000 per year minimum, requires sales call
- PitchBook API: more, requires sales call
- Harmonic.ai: enterprise pricing, requires sales call
- Forager.ai: enterprise pricing, requires sales call</p>
<p>Free-tier developer-investors and angels, exactly the people we built GitDealFlow for, get nothing. They watch the SEC EDGAR feed in their terminal. They scrape GitHub manually. They build their own signal stacks one Python script at a time.</p>
<h2>Then Google shipped A2A</h2>
<p>In April 2025, Google open-sourced a protocol called Agent2Agent. The spec is small:</p>
<ol><li>Publish a JSON file at `/.well-known/agent-card.json` describing what your agent can do.</li><li>Expose a JSON-RPC 2.0 endpoint that accepts `message/send` requests.</li><li>Other agents discover you, parse the card, and call your endpoint.</li></ol>
<p>That is it. There is no SaaS to license. There is no SDK to install if you do not want one. It is a contract between agents.</p>
<p>By April 2026 the protocol had passed 22,000 GitHub stars and 150+ organizations onboard. Microsoft, Salesforce, SAP, Cisco, ServiceNow, all in. Google's own Gemini Enterprise lets you register A2A agents in two clicks.</p>
<p>But none of those 150 organizations sells startup engineering signals. We do.</p>
<h2>The fix took a weekend</h2>
<p>So this weekend I shipped the GitDealFlow Signal Agent. Five skills, all free, no auth:</p>
<p>- `get_trending_startups`, top 20 across all sectors
- `search_startups_by_sector`, filtered by sector slug
- `get_startup_signal`, full profile for a named startup
- `get_signals_summary`, dataset metadata
- `get_methodology`, how we compute the signals</p>
<p>The AgentCard is at <a href="https://signals.gitdealflow.com/.well-known/agent-card.json">signals.gitdealflow.com/.well-known/agent-card.json</a>. The endpoint is at <a href="https://signals.gitdealflow.com/api/a2a">signals.gitdealflow.com/api/a2a</a>. The whole thing is roughly 250 lines of TypeScript inside our existing Next.js app.</p>
<p>You can call it right now from your terminal:</p>
<p>```
curl -X POST https://signals.gitdealflow.com/api/a2a \
  -H 'Content-Type: application/json' \
  -d '{
    &quot;jsonrpc&quot;: &quot;2.0&quot;,
    &quot;id&quot;: 1,
    &quot;method&quot;: &quot;message/send&quot;,
    &quot;params&quot;: {
      &quot;message&quot;: {
        &quot;role&quot;: &quot;user&quot;,
        &quot;parts&quot;: [{ &quot;kind&quot;: &quot;text&quot;, &quot;text&quot;: &quot;Show me trending startups&quot; }]
      }
    }
  }'
```</p>
<p>You will get back a Task with `status.state: &quot;completed&quot;` and 20 startups in the artifact data part. Each startup row carries commit velocity, contributor growth, signal classification, and a GitHub URL.</p>
<p>That same call works from any A2A-aware agent runtime: Google Agent Builder, LangChain, CrewAI, Mastra, Vercel AI SDK, Inkeep. Drop the AgentCard URL in. The agent figures out the rest.</p>
<p>If you want to skip the curl entirely, the <a href="/a2a-demo">interactive playground</a> lets you pick a skill, hit send, and watch the JSON come back.</p>
<h2>What this means in practice</h2>
<p>Three things that did not exist on Friday now exist on Monday.</p>
<p><strong>Your Claude can now do startup signal lookups inline.:</strong> No copy-paste from our website. No tab-switching. Tell Cursor &quot;find me a fintech startup with breakout commit velocity this quarter&quot; and Cursor calls our agent on its own.</p>
<p><strong>Other agent platforms can chain through us.:</strong> A diligence agent can call our agent, get a list of names, then call its own LinkedIn agent for headcount, then call its own SEC EDGAR agent for filings. We become one node in a directed graph of automated diligence.</p>
<p><strong>The dataset is now AI-cited.:</strong> ChatGPT and Perplexity will, over time, learn that there is a free A2A agent for VC signals. When a developer-investor asks a question in our domain, the model knows where to point.</p>
<h2>What the agent still cannot do</h2>
<p>The stub I shipped is read-only and synchronous. It does sync `message/send` returning a terminal Task, all five skills via text intent or structured data parts, CORS preflight, and JSON-RPC error codes. It does NOT yet do:</p>
<ol><li>Streaming via `message/stream`</li><li>Task persistence and `tasks/get`</li><li>Push notifications</li><li>Authenticated extended cards</li><li>Custom watchlists or per-user predictions</li></ol>
<p>That last bullet is the cliffhanger. The thing your AI can ask GitDealFlow today is &quot;who is trending in fintech.&quot; The thing you can ask <a href="/predict">/predict</a> is &quot;predict whether this specific startup will raise a Series A in the next 60 days.&quot; Different question, different answer. Both free at the moment, both heading toward a value ladder, but only one is callable by your AI yet.</p>
<p>When the prediction layer goes live as an A2A skill, that is when this gets interesting. Until then, the agent is the rung that catches your AI when it is looking for &quot;anything interesting in startup-land this week.&quot;</p>
<h2>Three things to do today</h2>
<p>If you are a developer-investor, an angel scout, or anyone running an AI workflow that touches startup data:</p>
<p><strong>Configure your AI runtime.:</strong> Drop the <a href="https://signals.gitdealflow.com/.well-known/agent-card.json">AgentCard URL</a> into Claude Code, Cursor, Windsurf, or your preferred agent. The five skills become available as tools. Two minutes.</p>
<p><strong>Bookmark the curl.:</strong> The example above works without any agent runtime. Fold it into your scripts.</p>
<p><strong>Cite us in your reports.:</strong> The free-tier license is &quot;Free for personal and editorial use, attribution required.&quot; Cite as `VC Deal Flow Signal (signals.gitdealflow.com), Q2 2026 data.` We have published the full methodology peer-reviewed at SSRN: <a href="https://ssrn.com/abstract=6606558">ssrn.com/abstract=6606558</a>.</p>
<h2>Why we are telling you instead of pitching you</h2>
<p>GitDealFlow's business model has always been: ship the free distribution magnet first, charge for the premium layer. Our 6 MCP tools are free forever. Our 5 A2A skills are free forever. The signals website is free. The CSV export is free. The methodology paper is free.</p>
<p>What is paid: <a href="/dashboard">Dashboard Beta</a> (€9.97/mo) and <a href="/insider">Insider Circle</a> (€97/mo). Those will earn their keep when there are 500 paying subscribers and we can start building the things only that audience needs.</p>
<p>Until then, every developer-investor who wires GitDealFlow into their Claude Code is one more node in a network where breakout startups cannot hide for 6-12 weeks anymore. That is worth more than $20,000 of API access.</p>
<p>Crunchbase API: $20,000 per year. GitDealFlow A2A: free, no signup. Your move.</p>]]></content:encoded>
      <pubDate>Sun, 26 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Every dev has invested in unicorns. They just don't know it.]]></title>
      <link>https://signals.gitdealflow.com/blog/receipts-launched</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/receipts-launched</guid>
      <description><![CDATA[I shipped Receipts at signals.gitdealflow.com/receipts. Paste your GitHub username, get a Scout Score from your starring history. The unicorns you starred before the news broke, Vercel at 200 stars, LangChain in week 2, OpenAI before the $157B round, are now worth points. Free, no login, no OAuth.]]></description>
      <content:encoded><![CDATA[<p>Saturday morning I shipped Receipts. Free tool, no login, eight seconds from paste to shareable card. The hook is one line: every dev has invested in unicorns, they just do not know it.</p>
<p>Here is what it does. You paste your GitHub username at <a href="https://signals.gitdealflow.com/receipts">signals.gitdealflow.com/receipts</a>. We fetch your public starring history. We cross-reference every starred repo against a curated database of ~75 validated unicorns, companies that hit a $1B+ valuation, raised a Series A or later, were acquired, or crossed 25K+ stars in the last five years. For each match we measure how many months early you starred. Twenty-four months early on a $1B-valuation company gets you the maximum 100 points. Five perfect calls is a Scout Score of 100.</p>
<p>That is the whole product. No login. No OAuth. No private repos read. Just public starring metadata that GitHub has been quietly broadcasting for fourteen years.</p>
<h2>Why backwards-looking</h2>
<p>We already shipped a forward-looking version. <a href="/predict">/predict</a> is a six-month prediction game where you call whether a startup will raise a Series A. It works. It has a leaderboard, an OG card, a welcome onboarding sequence, a public profile at /s/[handle]. But it has a virality ceiling I underestimated. Twitter does not share things that pay off in Q4.</p>
<p>Receipts is the inversion. You do not have to wait six months. The receipts already exist in your GitHub account. We just made them visible.</p>
<p>The first viral test was tj. TJ Holowaychuk's first three hundred starred repos contained two early calls: Deno (48.7 months early) and Tauri (42.5 months early). Scout Score: 22. Rank: Scout. The second test was sindresorhus. Five matched wins, four called early, top one OpenAI starred 24.4 months before the $157B valuation. Score: 36. Rank: Scout. Both cards rendered as 1200×630 PNGs and look exactly like the kind of thing developer Twitter shares to flex.</p>
<p>That is the whole bet. People share Spotify Wrapped. People share their Strava year. People share their Letterboxd top ten. There has never been a &quot;look at the unicorns I starred early&quot; card. Now there is.</p>
<h2>The scoring math</h2>
<p>I want to be precise about how the score works because the math is the trust layer. For each starred repo that matches a validated win:</p>
<p>```
months_early = (event_date - star_date) / 30.4375
if months_early &lt;= 0:
  points = 0   # You starred late. Sorry.
else:
  points = weight × min(months_early / 24, 1.0)
```</p>
<p>Weight scales with the event:
- Series A = 50
- Series B = 70
- Series C and beyond, or acquisition = 80 to 90
- $1B+ valuation or post-IPO = 100
- Mass-adoption milestone (25K+ stars without public funding) = 30 to 50</p>
<p>We dedupe by company, starring three Vercel repos counts as one win, taking your earliest star date. Top 5 wins are summed and normalized so five perfect early calls equals a Scout Score of 100. Rank ladder is shared with /predict: Curious → Scout → Sharp → Elite → Oracle.</p>
<p>The full validated-wins database is hardcoded JSON in the repo. It is not exhaustive. It is biased toward developer-tools, AI, and data/ops companies that have public GitHub presence. Closed-source unicorns are unrepresented. If your favorite unicorn is missing, your real Scout Score is higher than what we display. That is a known false-negative and we add to the list as we learn.</p>
<h2>What I left out</h2>
<p>The first version is intentionally under-engineered.</p>
<p><strong>No persistence.:</strong> Each receipt is computed on demand and cached for 24 hours via Vercel's CDN. We do not save your Scout Score to a database. We do not link your receipts to a /predict scout account unless you explicitly come over and predict. There is no &quot;claim your score&quot; gate, no email capture, no upsell on the result page. The CTA at the bottom is a soft link to /predict, not a modal.</p>
<p><strong>No AI commentary.:</strong> I considered adding an Anthropic API call to generate a one-paragraph &quot;your taste personality&quot; blurb. Decided against it: deterministic templates are faster, free, and reliable. The current commentary is built from category counts (&quot;you have a clear bias for AI/ML infrastructure&quot; / &quot;an operator's instinct for observability and workflow tools&quot;) with the top win surfaced by name. It will get smarter as the database grows.</p>
<p><strong>No comparison feature.:</strong> There is no &quot;compare with friend&quot; button. The viral loop is the permalink at /receipts/[username], every shared card has a stable URL that anyone can click and see the same result. Permalinks are the comparison surface; we do not need a UI for it.</p>
<p><strong>No fancy graphics.:</strong> The OG card is one big number, three early calls, a footer. No charts, no animations, no character avatars. The card has to read in 0.4 seconds in a Twitter feed. Anything more is friction.</p>
<h2>How the dev-loop ran</h2>
<p>Total ship time was about three hours from &quot;paste your GitHub username&quot; idea to live in production. The architecture leans hard on what was already built:</p>
<p>- **PocketBase** was not used. Receipts is stateless.
- **The OG image route** at /api/og/scout/[handle]/route.tsx was the template, same dark gradient, same monospace stat blocks, same rank colors. The Receipts card at /api/og/receipts/[username]/route.tsx is a 200-line variant.
- **The 5-email welcome sequence** at /lib/soap-opera-scout.ts already exists. Receipts does not trigger it. /predict does. We keep one funnel, not two.
- **The validated-wins database** is a static JSON file. About 75 entries. I curated by hand from memory, Crunchbase summaries, and a quick pass through the most-starred OSS repos from 2021-2025. It will be a moving target, every quarter has new unicorns and old &quot;wins&quot; get re-validated.</p>
<p>The trickiest part was the GitHub rate limit. Unauthenticated, the API gives you 60 requests per hour per IP. That is fine for a single user pulling their own history (one to three paginated calls). It dies the second the launch tweet goes out. The fix was a fine-grained PAT with zero scopes, even an empty-permission token bumps the rate limit to 5,000 per hour. Plus an in-memory cache per Vercel Function instance. Plus the 24-hour CDN cache via stale-while-revalidate. Together those handle a viral spike without ceremony.</p>
<h2>The bigger play</h2>
<p>The reason Receipts exists is to feed /predict. The conversion path is:</p>
<ol><li>Receipt card goes viral on Twitter.</li><li>Friend sees it, clicks, enters their own username.</li><li>Friend gets a Scout Score, learns the ranks (Curious → Scout → Sharp → Elite → Oracle).</li><li>The result page CTA links to /predict.</li><li>They make a forward-looking call. The Scout Game starts. The welcome sequence kicks off.</li></ol>
<p>Receipts is a free top-of-funnel for a paid product. Same brand, same vocabulary, same ranks, different timing. /predict resolves in six months. Receipts resolves in eight seconds. We can finally tell people why they should care today.</p>
<h2>What I want from you</h2>
<p>If you have a public GitHub account, paste your username at <a href="https://signals.gitdealflow.com/receipts">signals.gitdealflow.com/receipts</a>. Three things help:</p>
<ol><li>**Tweet the card.** Cards have stable permalinks at /receipts/[your-username]. Anyone who clicks gets their own card. The viral loop is the whole ROI.</li><li>**Tell me what is missing.** If you starred a unicorn we do not have in the database, the score under-counts you. Reply with the org and the validation event and I will add it.</li><li>**Now go predict.** The receipts are backwards. The Scout game is forwards. <a href="/predict">/predict</a> is where the points keep flowing.</li></ol>
<p>The receipts already exist. I just made them visible.</p>]]></content:encoded>
      <pubDate>Sun, 26 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Free Scout Score badges: shields.io for GitHub investing taste.]]></title>
      <link>https://signals.gitdealflow.com/blog/scout-badge-launched</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/scout-badge-launched</guid>
      <description><![CDATA[I shipped two free SVG badges for any GitHub README. One renders your live Scout Score (0-100) from your starring history. The other renders the live commit-momentum tier of any tracked repo. Same shields.io look as Codecov / WakaTime, auto-updates, no signup, no telemetry.]]></description>
      <content:encoded><![CDATA[<p>Two endpoints went live tonight. Both return SVG. Both are free. Both are autonomous traffic compounders.</p>
<p>```markdown
<a href="https://signals.gitdealflow.com/api/badge/scout/torvalds/svg">![Scout Score</a>](https://signals.gitdealflow.com/badge-builder)
<a href="https://signals.gitdealflow.com/api/badge/momentum/mlflow/mlflow/svg">![Commit Momentum</a>](https://signals.gitdealflow.com/badge-builder)
```</p>
<p>That is the whole product. Paste those lines into any README, replace the username and the org/repo, you have a live shields.io-style badge that auto-updates whenever your starring history grows or the tracked repo's commit velocity moves a tier. Builder UI with copy-paste markdown / HTML / BBCode at <a href="https://signals.gitdealflow.com/badge-builder">signals.gitdealflow.com/badge-builder</a>.</p>
<h2>Why this is the right move right now</h2>
<p>We have an MCP server, a paper on SSRN, a Wikidata entity, twenty-plus blog posts, three working AI-discovery surfaces (llms.txt, agent-card.json, ai-plugin.json), and dozens of channel listings. Most of those are read-once. Someone discovers us, decides whether to subscribe, and either bookmarks or moves on.</p>
<p>A README badge inverts that. The badge gets pasted once and then renders every time the host page loads, every issue triage, every PR review, every visitor to the repo. Each render is one impression for our domain through GitHub's camo proxy. The badge is the gift that keeps giving.</p>
<p>Codecov did this. WakaTime did this. GitHub Stats did this. Shields.io did this. The pattern is so well-established that a casual reader scanning a README does not even register the badge as marketing, it reads as a normal artifact of &quot;this maintainer takes their stack seriously.&quot;</p>
<h2>What it shows</h2>
<p>The Scout Score badge renders the same 0-100 number that already exists at <a href="https://signals.gitdealflow.com/receipts">/receipts/{username}</a>, computed live from a user's public GitHub starring history vs. the validated-wins database. The color tracks the rank: curious (teal) → scout (sky) → sharp (purple) → elite (amber) → oracle (rose). The user clicks the badge, lands on the builder, can pull their own snippet in fifteen seconds.</p>
<p>The Commit Momentum badge renders the live commit-velocity tier for any tracked GitHub org. Tiers are deterministic from the 14-day velocity change versus the prior 14-day window:</p>
<p>```
breakout : &gt;= +200%
hot      : &gt;= +50%
warming  : &gt;= -30%
cold     : &lt;  -30%
```</p>
<p>Untracked orgs render a neutral &quot;untracked&quot; pill so a maintainer can paste the badge before we have indexed their repo without breaking the README. When we add the org to the next weekly crawl, the badge starts rendering the real tier, the maintainer does not have to do anything.</p>
<h2>The technical guardrails</h2>
<p>A README badge has exactly one job: never break. A broken-image icon in a README gets the badge removed within hours, and the maintainer never trusts the source again. Three rules:</p>
<ol><li>**Always return 200.** Bad input, rate limit, GitHub timeout, internal error, all paths render a neutral gray &quot;pending&quot; SVG. Never a 4xx, never a 5xx, never a JSON error.</li><li>**Aggressive cache.** `Cache-Control: public, max-age=300, s-maxage=86400, stale-while-revalidate=604800`. Browser holds 5 min, CDN holds 24h, falls back to stale up to 7 days while it re-fetches. GitHub's camo proxy respects the ETag (`{username}:{score}:{rank}`), so revalidation is one HTTP HEAD per camo region per hour.</li><li>**No new infra.** Both endpoints reuse the existing Scout Score primitives, the existing in-memory star cache, the existing per-IP rate limiter, the existing rank colors. Zero new dependencies. Zero new env vars. Zero new hosting cost.</li></ol>
<p>The whole thing is two route handlers, one shared SVG generator, and one client component for the builder UI. Total diff is under 500 lines. It deploys with a single `vercel build &amp;&amp; vercel deploy --prebuilt --prod`.</p>
<h2>How the distribution loop closes</h2>
<p>The badge is now linked from four surfaces inside our own product:</p>
<ol><li>**The MCP server README on npm + Glama + GitHub.** Anyone landing on our most-discovered surface sees the badge in the wild and learns it exists.</li><li>**The /receipts result page.** Right after a user gets their Scout Score, they see &quot;show off your taste&quot; with the markdown ready to copy.</li><li>**The /s/[handle] scout profile page.** Small footer link, anyone visiting a public scout profile can grab the same badge.</li><li>**The /developers page and /badge-builder UI.** High-intent dev visitors discover both endpoints with copy-paste snippets.</li></ol>
<p>Plus the OpenAPI spec, llms.txt, llms-full.txt, every AI assistant that reads our metadata will surface the badge endpoints when asked about our API.</p>
<p>The bet is that once a few high-signal accounts paste it (a couple of MCP-curious devs, one or two scouts on the leaderboard), the badge spreads on its own. Every render after that is free traffic that compounds with the leaderboard growth.</p>
<h2>What I want from you</h2>
<p>If you have a GitHub profile and you want to flex your taste, paste this into your profile README and replace `YOUR-USERNAME`:</p>
<p>```markdown
<a href="https://signals.gitdealflow.com/api/badge/scout/YOUR-USERNAME/svg">![Scout Score</a>](https://signals.gitdealflow.com/badge-builder)
```</p>
<p>If you maintain a tracked OSS repo and want to show your commit momentum, replace `ORG/REPO`:</p>
<p>```markdown
<a href="https://signals.gitdealflow.com/api/badge/momentum/ORG/REPO/svg">![Commit Momentum</a>](https://signals.gitdealflow.com/badge-builder)
```</p>
<p>Or just open <a href="https://signals.gitdealflow.com/badge-builder">/badge-builder</a>, paste a handle, click copy. Three seconds. Free forever.</p>]]></content:encoded>
      <pubDate>Sun, 26 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[I cut my MCP server from 8 tools to 5 and the hallucinations stopped]]></title>
      <link>https://signals.gitdealflow.com/blog/mcp-server-tool-count-war-story</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/mcp-server-tool-count-war-story</guid>
      <description><![CDATA[Three weeks of tool-count post-mortem on @gitdealflow/mcp-signal. Why REST endpoints aren't user intents, why two of my tool names were costing me selection accuracy, and the data on what changed.]]></description>
      <content:encoded><![CDATA[<p>I shipped <a href="https://www.npmjs.com/package/@gitdealflow/mcp-signal">@gitdealflow/mcp-signal</a> in two hours. Eight tools, mirrored one-to-one off our REST API. Felt clean. Looked clean. The first time I plugged it into Claude and asked &quot;what's the trending startup in fintech this week,&quot; it called `get_startup` with a `sector` parameter that doesn't exist, hallucinated a result, and confidently quoted me numbers nobody had ever calculated.</p>
<p>I wasn't building an MCP server. I was building a very expensive random-number generator with a JSON wrapper.</p>
<p>The next three weeks were a tool-count post-mortem. The end state was 5 tools, two of them renamed for verb-clarity, and a selection accuracy that went from &quot;wrong about a third of the time&quot; to &quot;I genuinely cannot remember the last time it picked the wrong one.&quot; Here is what actually mattered.</p>
<h2>The 8 tools that didn't work</h2>
<p>The starting menu, copied straight from the REST routes:</p>
<p>```
list_startups
get_startup
list_signals
get_signal
get_methodology
get_trending
list_sectors
get_sector
```</p>
<p>Reasonable on paper. Each tool described a real capability. The schemas were valid. The descriptions read like decent docstrings. I was even proud of how cleanly the surface mapped onto the API.</p>
<p>The problem became obvious the first time I watched a real conversation. The user asked something agentic, &quot;show me the top fintech startups this week, and tell me what makes them interesting.&quot; A single intent, a single ranked list. Claude did this:</p>
<ol><li>Called `list_sectors` (probably to confirm &quot;fintech&quot; is a valid sector)</li><li>Called `list_startups` with a `sector` parameter (does not exist on `list_startups`, schema rejected it)</li><li>Retried with a different parameter shape</li><li>Eventually gave up and called `get_trending`</li><li>Made up a `top_n` parameter that does not exist either</li><li>Returned a &quot;this is what is trending&quot; answer that was actually four random startups from cache</li></ol>
<p>Five tool calls for a single intent. Three of them with hallucinated parameters. Zero of them returned the data the user actually wanted.</p>
<p>This is the part nobody talks about when they say &quot;MCP just works.&quot; It works in demos because demo prompts map cleanly to one tool. It stops working the moment a user asks something composite, which is most user prompts.</p>
<h2>What the model is actually doing</h2>
<p>I had been thinking about MCP tools as endpoints. The model thinks about them as items on a menu it has to read every single turn.</p>
<p>Eight tools means eight schemas in context. Each schema includes a description, a parameter list, parameter types, parameter descriptions, return type, return description. Even with terse docstrings, that runs ~600 tokens per tool. Eight tools ≈ 5,000 tokens of menu before the user has said anything.</p>
<p>Worse, the model has to hold all eight in working memory while it picks one. The picking process is essentially a vector-similarity beauty contest: the user prompt's embedding against each tool description's embedding. If two tools have descriptions that share half their vocabulary, `list_startups` and `get_startup`, say, both heavy on the word &quot;startup&quot;, the model's confidence between them collapses to something close to a coin flip.</p>
<p>Most &quot;the AI hallucinated a tool call&quot; stories I have heard in the last quarter are this exact failure. Not a model failure. A menu design failure.</p>
<h2>The cuts</h2>
<p>Three tools got dropped in the first pass. Two more got renamed.</p>
<p><strong>`list_startups` and `get_startup` were collapsed.:</strong> The model was confusing them on every other turn. The tell: when I logged the model's reasoning, it would describe what it wanted as &quot;a list-style get of startups in fintech&quot;, which is actually `list_startups(sector=&quot;fintech&quot;)`. But it kept calling `get_startup` with a `sector` parameter, because the names were too close.</p>
<p>I killed `get_startup` entirely. If you want a single startup, you call the list tool with a filter and `limit=1`. The single-resource endpoint had been costing me selection accuracy without buying any real capability.</p>
<p><strong>`list_sectors` and `get_sector` went the same way.:</strong> Almost nobody, including the model, wanted a single sector. They wanted a list to pick from, which used to be `list_sectors`'s job. I rolled both into a single tool I named `search_startups_by_sector`, a verb-noun-prepositional-phrase shape that the model parses extremely cleanly. &quot;Find me fintech startups&quot; → unambiguous match.</p>
<p><strong>`list_signals` got renamed to `get_startup_signal`.:</strong> This was the subtle one. The user almost never says &quot;signals&quot; in a prompt, they say &quot;what's the engineering activity look like for X&quot; or &quot;is this team building.&quot; The word &quot;signals&quot; is internal jargon. The rename made the model start picking the right tool on prompts that did not even contain the word signal, because it parsed &quot;startup&quot; from context and matched on that.</p>
<p><strong>`get_trending` got renamed to `get_trending_startups`.:</strong> Same idea. Verb-adjective-noun where the noun is a word the user actually said is a much stronger lock than verb-adjective alone.</p>
<h2>The 5 that work</h2>
<p>```
get_trending_startups
search_startups_by_sector
get_startup_signal
get_signals_summary
get_methodology
```</p>
<p>Two things worth calling out beyond the renames.</p>
<p><strong>`get_signals_summary` is a new tool.:</strong> It does not have a 1:1 REST endpoint. It exists because users kept asking &quot;give me a one-paragraph summary of what's interesting this week&quot; and the model kept stitching together three calls to fake it. I built the summary tool. The model now makes one call.</p>
<p>That last point is the heuristic I would give anyone shipping an MCP server: look at the actual conversational intents your users have, and design one tool per intent. Resources are an implementation detail. Intents are what the model is selecting against.</p>
<p><strong>Verb-noun, not noun-noun.:</strong> Even the kept tools got their names re-checked against this rule. `get_methodology` survived because users do say &quot;methodology&quot;, but if I noticed selection drift I would rename it `describe_signal_methodology` to anchor on the verb.</p>
<h2>The data, three weeks later</h2>
<p>I have logging on every tool call. Before the cuts, on a sample of 200 real prompts:</p>
<p>- 132 / 200 (66%) ended in a correct tool selection on the first call
- 68 / 200 (34%) involved at least one hallucinated parameter or wrong-tool selection
- Average tool calls per user intent: 2.4</p>
<p>After the cuts, on the same kinds of prompts:</p>
<p>- 197 / 200 (98.5%) ended in a correct tool selection on the first call
- 3 / 200 (1.5%) involved a wrong-tool selection (all three were obscure edge cases)
- Average tool calls per user intent: 1.1</p>
<p>The token cost of menu inflation went from ~5,000 input tokens per turn to ~2,800. Selection accuracy effectively saturated. And, this is the part I underestimated, the model's *latency on the first token* dropped noticeably, because it was no longer chewing through eight schemas before picking one.</p>
<h2>The thing I would tell past me</h2>
<p>If I could go back to the morning I shipped the first version, I would tell myself two things.</p>
<p>First: your tool count is a liability, not an asset. Every tool you add costs the model reasoning, costs the user latency, and costs you accuracy. Every tool needs to earn its place by mapping to a distinct user intent that no other tool maps to.</p>
<p>Second: REST API endpoints are not user intents. The clean mental model is: &quot;what would a user say to express this need,&quot; not &quot;what HTTP route serves this resource.&quot; The mapping is rarely 1:1. Most APIs have more endpoints than they have distinct user intents, ours had eight endpoints and five intents, and shipping the extra three as MCP tools is just paying the menu tax for nothing.</p>
<p>I am at five tools and I think four of them are load-bearing. The fifth (`get_methodology`) only fires maybe 1 in 50 conversations, and I am watching it. If selection accuracy on the other four starts degrading, that is the next cut.</p>
<p>The MCP spec lets you ship as many tools as you want. The model does not reward you for shipping more.</p>
<h2>How to use this</h2>
<p>If you are shipping an MCP server, the audit is fast:</p>
<ol><li>Log every tool call from a representative week of real conversations.</li><li>Cluster the user prompts by intent in plain English (ignore the tool the model picked).</li><li>Count distinct intents.</li><li>If your tool count exceeds your intent count, you have menu inflation. Cut to match.</li></ol>
<p>The repo for the five-tool version is at <a href="https://github.com/kindrat86/mcp-deal-flow-signal">github.com/kindrat86/mcp-deal-flow-signal</a>. The schemas, the descriptions, and the changelog are all there.</p>
<p>If you have shipped an MCP server, what is your tool count and how did you arrive at it? Public reporting on this trade-off is surprisingly thin and I am collecting examples.</p>]]></content:encoded>
      <pubDate>Sat, 25 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[47 Alternative Data Sources for Angel Investors in 2026]]></title>
      <link>https://signals.gitdealflow.com/blog/47-alternative-data-sources-angel-investors-2026</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/47-alternative-data-sources-angel-investors-2026</guid>
      <description><![CDATA[Most angel investors check 3 sources. Here are 47 signals that catch startups 6-12 weeks before Crunchbase, from GitHub velocity to SEC Form D filings.]]></description>
      <content:encoded><![CDATA[<p>Most angel investors only check 3 data sources: Crunchbase, LinkedIn, and the pitch deck. The problem: those are all lagging. By the time a startup shows up there, the signal is already priced into the round.</p>
<p>Here are 47 alternative data sources that surface breakout startups 6-12 weeks before they hit a VC database. We publish our methodology at <a href="https://ssrn.com/abstract=6606558">ssrn.com/abstract=6606558</a>.</p>
<p>Table of contents</p>
<ol><li>Public code signals (6)</li><li>Hiring and team signals (5)</li><li>Product telemetry (5)</li><li>Infrastructure buildouts (5)</li><li>Search and attention (5)</li><li>Funding leading indicators (4)</li><li>Community signals (3)</li><li>Regulatory and IP (4)</li><li>Open-data scientific (5)</li><li>Niche commercial APIs (5)</li></ol>
<h2>1. Public code signals</h2>
<p>The earliest signal of startup momentum is usually in public code. A team shipping fast leaves a timestamped trail that is hard to fake and easy to query.</p>
<ol><li>GitHub commit velocity, contributors, stars. Track the rate of change in 14-day commit counts across a company's public repositories, not absolute output. A doubling in two weeks typically means a hiring burst or a product sprint. Contributor graph growth confirms headcount expansion; star velocity on a founder's personal repo often flags dev-tool launches weeks before Product Hunt. Access: free REST and GraphQL APIs with 5,000 requests per hour for authenticated users. Lead time: 6-12 weeks. Start at <a href="https://docs.github.com/en/rest">docs.github.com/en/rest</a>.</li></ol>
<ol><li>GitLab public project activity. GitLab hosts a smaller but non-overlapping population, especially European dev-tool and infra companies. The activity API returns push events, issue creation, and merge requests with timestamps. Access: free API, public-project scope is open. Lead time: 4-10 weeks. Start at <a href="https://docs.gitlab.com/ee/api/events.html">docs.gitlab.com/ee/api/events.html</a>.</li></ol>
<ol><li>npm package downloads. Weekly download counts for a startup's published packages correlate with developer adoption. A 3x week-over-week jump on a niche package often precedes a Series A. Access: free downloads API. Lead time: 4-8 weeks. Start at <a href="https://github.com/npm/registry/blob/main/docs/download-counts.md">github.com/npm/registry, download counts</a>.</li></ol>
<ol><li>PyPI download statistics. Python equivalents via BigQuery's public dataset, which exposes per-package daily downloads since 2016. Useful for ML, data, and scientific-tooling startups. Access: free via Google BigQuery (first 1TB per month is free). Lead time: 4-8 weeks. Start at <a href="https://pypistats.org">pypistats.org</a>.</li></ol>
<ol><li>Docker Hub image pulls. Public image pull counts reveal production adoption. A containerized product crossing 10k monthly pulls is usually past proof-of-concept. Access: pull counts visible on the public page; batch via the unofficial API. Lead time: 3-6 weeks. Start at <a href="https://hub.docker.com">hub.docker.com</a>.</li></ol>
<ol><li>Release cadence on GitHub Releases. The gap between tagged releases is a clean proxy for shipping velocity. A move from monthly to weekly releases usually means a team scale-up. Access: free, part of the GitHub REST API. Lead time: 4-8 weeks. Start at <a href="https://docs.github.com/en/rest/releases">docs.github.com/en/rest/releases</a>.</li></ol>
<h2>2. Hiring and team signals</h2>
<p>Headcount moves before revenue, and hiring intent shows up in public job boards weeks before offers are signed.</p>
<ol><li>LinkedIn employee count over time. Track the delta in headcount and function mix. A jump in senior engineering hires is a classic pre-Series-A signal. Access: free via company pages; scraping is against ToS, so use the LinkedIn Sales Navigator API or a licensed provider. Lead time: 4-8 weeks. Start at <a href="https://www.linkedin.com/">linkedin.com</a>.</li></ol>
<ol><li>Hacker News Who Is Hiring threads. Monthly thread where YC and non-YC startups post roles directly. A company appearing for the first time, especially with multiple roles, signals fresh capital. Access: free, use the HN Algolia API. Lead time: 3-6 weeks. Start at <a href="https://hn.algolia.com/api">hn.algolia.com/api</a>.</li></ol>
<ol><li>AngelList Talent job postings. Wellfound (formerly AngelList Talent) still surfaces early-stage roles with salary and equity ranges. The equity band often hints at the company's stage more clearly than the pitch deck. Access: free browsing, paid API. Lead time: 3-6 weeks. Start at <a href="https://wellfound.com">wellfound.com</a>.</li></ol>
<ol><li>Gem.com outbound hiring signals. Gem aggregates anonymized recruiter outbound activity. A spike in recruiter outreach for a company or from its recruiters is a compound signal: hiring, capital, and sector heat. Access: paid API, designed for recruiting teams. Lead time: 4-6 weeks. Start at <a href="https://gem.com">gem.com</a>.</li></ol>
<ol><li>Indeed and general job-board velocity. Indeed's free search surfaces the same roles across the long tail of job sites. The week-over-week change in posted-role count per company is noisy but useful for cross-checks. Access: free search, paid API. Lead time: 2-4 weeks. Start at <a href="https://indeed.com">indeed.com</a>.</li></ol>
<h2>3. Product telemetry</h2>
<p>Public product surfaces leak more than teams realize. Launch platforms, tech stacks, and traffic estimators all give early reads on adoption.</p>
<ol><li>Product Hunt launches. Even failed PH launches tell you the founder has shipped and is running distribution. The velocity of a maker's past PH launches often indicates pace. Access: free GraphQL API. Lead time: 2-6 weeks. Start at <a href="https://api.producthunt.com/v2/docs">api.producthunt.com</a>.</li></ol>
<ol><li>BuiltWith tech stack changes. BuiltWith logs the JavaScript, CMS, analytics, and payment stack of every crawled domain over time. A move from Stripe to a custom billing stack usually means enterprise pivot. Access: free for single-domain lookups, paid API for bulk. Lead time: 4-8 weeks. Start at <a href="https://builtwith.com">builtwith.com</a>.</li></ol>
<ol><li>Similarweb traffic estimates. Directional traffic data that is noisy below 5k monthly visits but solid above. A 3x month-over-month lift on a product subdomain is almost always real growth. Access: free tier with limits, paid API. Lead time: 4-6 weeks. Start at <a href="https://similarweb.com">similarweb.com</a>.</li></ol>
<ol><li>Wappalyzer tech detection. Browser extension and API that fingerprints frontend and backend stacks. Useful to spot when a startup migrates from a no-code stack to a custom build, which usually means Series A traction. Access: free extension, paid API. Lead time: 4-8 weeks. Start at <a href="https://wappalyzer.com">wappalyzer.com</a>.</li></ol>
<ol><li>App Annie, now data.ai, rank shifts. Mobile-app App Store and Play Store rank data. A consumer app breaking into a category top-50 is visible 1-2 months before mainstream coverage. Access: free web views, paid API. Lead time: 4-6 weeks. Start at <a href="https://data.ai">data.ai</a>.</li></ol>
<h2>4. Infrastructure buildouts</h2>
<p>The pipes a startup builds are often the clearest sign of ambition. Infrastructure commitments precede revenue by months.</p>
<ol><li>AWS IP blocks and service footprint. AWS publishes its IP ranges as JSON. Cross-referencing against Passive DNS reveals when a company spins up new regions, which is a capital-commitment signal. Access: free JSON feed, Passive DNS tools are freemium. Lead time: 6-10 weeks. Start at <a href="https://ip-ranges.amazonaws.com/ip-ranges.json">ip-ranges.amazonaws.com/ip-ranges.json</a>.</li></ol>
<ol><li>DNS record changes. SecurityTrails, DNSDumpster, and RiskIQ all let you diff a company's DNS records over time. New subdomains like api.v2, eu.company.com, or enterprise.company.com all hint at product direction. Access: freemium across providers. Lead time: 4-8 weeks. Start at <a href="https://securitytrails.com">securitytrails.com</a>.</li></ol>
<ol><li>SSL certificate transparency logs. Every TLS cert issued is logged publicly and searchable. New cert for a stealth domain owned by a founder is a classic stealth-launch signal. Access: free via crt.sh. Lead time: 6-12 weeks. Start at <a href="https://crt.sh">crt.sh</a>.</li></ol>
<ol><li>Cloudflare RADAR. Free dashboard of DNS, HTTP, and attack trends at the network level. The per-domain traffic tab gives directional traffic and is harder to game than Similarweb. Access: free. Lead time: 2-6 weeks. Start at <a href="https://radar.cloudflare.com">radar.cloudflare.com</a>.</li></ol>
<ol><li>HTTP Archive. Bi-monthly crawl of the top-million sites with full waterfall data. Useful to spot when a startup cleans up its frontend or adds a CDN, both common pre-launch moves. Access: free, data in BigQuery public datasets. Lead time: 4-8 weeks. Start at <a href="https://httparchive.org">httparchive.org</a>.</li></ol>
<h2>5. Search and attention</h2>
<p>Attention is the oldest leading indicator. Search volume, forum momentum, and preprint citations all show demand before revenue.</p>
<ol><li>Google Trends. Free normalized search-volume data for any term or brand. A rising 90-day chart for a startup's name, ahead of its sector, usually means the founder has figured out distribution. Access: free web UI, unofficial APIs. Lead time: 2-6 weeks. Start at <a href="https://trends.google.com">trends.google.com</a>.</li></ol>
<ol><li>Reddit post velocity. Reddit's search API exposes post and comment counts per subreddit per keyword over time. A startup getting organic r/selfhosted or r/ClaudeAI mentions is often 1-2 months ahead of its PR cycle. Access: free API with rate limits. Lead time: 3-6 weeks. Start at <a href="https://reddit.com/dev/api">reddit.com/dev/api</a>.</li></ol>
<ol><li>Hacker News score curves. The HN Algolia API returns every submission with its final score and comment count. A Show HN that crosses 200 points in a niche is a durable signal, especially for dev-tool companies. Access: free. Lead time: 2-8 weeks. Start at <a href="https://hn.algolia.com/api">hn.algolia.com/api</a>.</li></ol>
<ol><li>arXiv citation lift. Semantic Scholar and OpenAlex both expose citation counts per paper per month. A preprint that triples citations in 30 days often accompanies a spinout. Access: free via Semantic Scholar API. Lead time: 8-16 weeks. Start at <a href="https://api.semanticscholar.org">api.semanticscholar.org</a>.</li></ol>
<ol><li>Substack subscriber growth. Founders who run newsletters expose subscriber counts on their public pages. A jump from 5k to 20k over a quarter is a distribution signal that usually precedes a product launch. Access: free web scraping, ToS permitting. Lead time: 4-8 weeks. Start at <a href="https://substack.com">substack.com</a>.</li></ol>
<h2>6. Funding leading indicators</h2>
<p>Regulatory filings are the closest thing to a legally required fundraise announcement, and they are free.</p>
<ol><li>SEC Form D filings. Every US private offering must file Form D within 15 days of first sale. The free EDGAR full-text search lets you query by issuer, promoter, or amount. Access: free. Lead time: 4-8 weeks before press coverage. Start at <a href="https://efts.sec.gov/LATEST/search-index">efts.sec.gov/LATEST/search-index</a>.</li></ol>
<ol><li>EDGAR full-text search. Beyond Form D, EDGAR indexes S-1, 10-K, and 8-K filings from public companies that often disclose acquisitions, investments, or partnerships with private startups. Access: free. Lead time: 2-6 weeks. Start at <a href="https://efts.sec.gov/LATEST/search-index">efts.sec.gov/LATEST/search-index</a>.</li></ol>
<ol><li>Companies House UK. Free register of every UK company, with director changes, charges, and confirmation statements. A new Series A usually shows up as a charge registered against new preference shares. Access: free API. Lead time: 2-4 weeks. Start at <a href="https://find-and-update.company-information.service.gov.uk">find-and-update.company-information.service.gov.uk</a>.</li></ol>
<ol><li>French INPI. France's national company register, equivalent to Companies House, with capital increases and articles of incorporation. Free access via data.inpi.fr. Access: free. Lead time: 2-4 weeks. Start at <a href="https://data.inpi.fr">data.inpi.fr</a>.</li></ol>
<h2>7. Community signals</h2>
<p>A startup's community growth rate is sometimes the only metric the founder cares about, and it is often publicly visible.</p>
<ol><li>Discord server growth. Public Discord servers expose member counts via server invites. Track weekly counts; a 2x jump in a month often signals a launch. Access: free, use the Discord API with a bot token. Lead time: 2-6 weeks. Start at <a href="https://discord.com/developers/docs">discord.com/developers/docs</a>.</li></ol>
<ol><li>Slack community join rate. Many dev-tool startups run Slack communities with public invite links. Some expose member counts directly; for others, a simple periodic join-ping tracks the count. Access: free via the workspace's own public signals. Lead time: 3-6 weeks. Start at <a href="https://slack.com/community">slack.com/community</a>.</li></ol>
<ol><li>Telegram subscriber velocity. Public Telegram channels show subscriber counts directly. Track the weekly delta; in crypto and consumer categories, a rising channel often precedes a token or app launch. Access: free via the Telegram Bot API. Lead time: 2-4 weeks. Start at <a href="https://core.telegram.org/bots/api">core.telegram.org/bots/api</a>.</li></ol>
<h2>8. Regulatory and IP</h2>
<p>Patent, clinical, and regulatory filings are the most underrated signal class. They are legally dated, structured, and free.</p>
<ol><li>USPTO patent filings. Full-text search of every US patent and application, including inventor and assignee. A startup assigning three patents in six months is either pre-IPO or pre-acquisition. Access: free via PatentsView API. Lead time: 12-24 weeks. Start at <a href="https://patentsview.org/apis">patentsview.org/apis</a>.</li></ol>
<ol><li>FDA 510k clearances. Medical-device clearances are publicly searchable. A 510k clearance is often the trigger for a Series A or B in medtech. Access: free. Lead time: 4-12 weeks. Start at <a href="https://fda.gov/medical-devices/510k-clearances">fda.gov/medical-devices/510k-clearances</a>.</li></ol>
<ol><li>ClinicalTrials.gov. Every US-regulated clinical trial is registered with sponsor, phase, and primary endpoint. A new Phase 2 trial from a private biotech is a capital-event leading indicator. Access: free API. Lead time: 8-16 weeks. Start at <a href="https://clinicaltrials.gov/data-api">clinicaltrials.gov/data-api</a>.</li></ol>
<ol><li>EU MDR database. EUDAMED tracks every medical device in the EU market. Registration of a new device by a private company often precedes EU commercial launch by a quarter. Access: free web search. Lead time: 6-12 weeks. Start at <a href="https://ec.europa.eu/tools/eudamed">ec.europa.eu/tools/eudamed</a>.</li></ol>
<h2>9. Open-data scientific</h2>
<p>Science happens in public now. Preprints, dataset releases, and model uploads all reveal what research-heavy startups are actually building.</p>
<ol><li>SSRN papers. Social Science Research Network hosts working papers across economics, finance, and legal research. A startup's founding team publishing on SSRN often indicates the academic anchor for the thesis. Access: free. Lead time: 12-24 weeks. Start at <a href="https://ssrn.com">ssrn.com</a>.</li></ol>
<ol><li>bioRxiv preprints. The main biology preprint server. A startup's scientific advisors publishing a preprint in the company's domain is often the first public trace of the thesis. Access: free API. Lead time: 12-26 weeks. Start at <a href="https://api.biorxiv.org">api.biorxiv.org</a>.</li></ol>
<ol><li>medRxiv preprints. The clinical medicine counterpart to bioRxiv. Useful for surfacing digital-health and clinical-decision startups before they incorporate. Access: free. Lead time: 12-26 weeks. Start at <a href="https://medrxiv.org">medrxiv.org</a>.</li></ol>
<ol><li>OpenAlex index. Open-data replacement for Microsoft Academic Graph, with every scholarly work indexed and citation graphs exposed. Useful for tracking a founder's publication trajectory. Access: free API. Lead time: 12-24 weeks. Start at <a href="https://openalex.org">openalex.org</a>.</li></ol>
<ol><li>Hugging Face models and datasets. Every public model or dataset on HF has a timestamped upload and download history. A startup releasing a flagship open model is usually doing distribution before a product launch. Access: free API. Lead time: 4-12 weeks. Start at <a href="https://huggingface.co/docs/hub/api">huggingface.co/docs/hub/api</a>.</li></ol>
<h2>10. Niche commercial APIs</h2>
<p>These are paid or semi-paid services that aggregate several of the above into investor-ready feeds. Only worth the cost once you are sourcing at scale.</p>
<ol><li>Crunchbase API. The standard, though lagging. Still useful as a baseline and for entity disambiguation across other sources. Access: paid API, limited free tier via the website. Lead time: 0-2 weeks. Start at <a href="https://data.crunchbase.com/docs">data.crunchbase.com/docs</a>.</li></ol>
<ol><li>Specter API. Specter aggregates employee count, web traffic, funding, and tech-stack signals into a single company record. Good for mid-stage sourcing. Access: paid. Lead time: 2-6 weeks. Start at <a href="https://tryspecter.com">tryspecter.com</a>.</li></ol>
<ol><li>Synaptic. Focused on consumer and mobile, pulls app-store data, web traffic, and paid-marketing spend into a unified feed. Access: paid, enterprise pricing. Lead time: 2-6 weeks. Start at <a href="https://synaptic.com">synaptic.com</a>.</li></ol>
<ol><li>Predictleads. Monitors company websites for job postings, technology changes, and press releases. Designed for sales teams but useful as a funding-intent feed. Access: paid API. Lead time: 2-4 weeks. Start at <a href="https://predictleads.com">predictleads.com</a>.</li></ol>
<ol><li>4Degrees. CRM-native signal layer that pipes relationship and news data into a deal-flow pipeline. More useful to partners managing a portfolio than to solo angels. Access: paid. Lead time: 2-4 weeks. Start at <a href="https://4degrees.ai">4degrees.ai</a>.</li></ol>
<h2>Putting it together</h2>
<p>No solo angel will monitor all 47 in real time. The point is to pick a subset that matches your sector and stage.</p>
<p>If you invest in dev tools, AI, data infrastructure, or developer-facing SaaS, categories 1 through 4 cover most of the public-code surface. That is the subset <a href="/">GitDealFlow aggregates</a> across 15 sectors, scored and ranked weekly. The <a href="/predict">Prediction Game</a> turns it into a public track record you can share.</p>
<p>For biotech, medtech, and research-heavy verticals, categories 8 and 9 give the longest lead time. For consumer, categories 3 and 5 move fastest.</p>
<p>Our methodology for using the first four categories to predict fundraises is published on SSRN at <a href="https://ssrn.com/abstract=6606558">ssrn.com/abstract=6606558</a>, with the underlying dataset open on Hugging Face. The single most important rule: measure change from a company's own baseline, not absolute output. That filters out the docs sprints, the CI noise, and the popularity that does not convert into product.</p>]]></content:encoded>
      <pubDate>Wed, 22 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[I Tracked 350+ Startup GitHub Orgs for Six Months. Here's What Predicts a Series A.]]></title>
      <link>https://signals.gitdealflow.com/blog/i-tracked-369-startup-github-orgs-six-months</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/i-tracked-369-startup-github-orgs-six-months</guid>
      <description><![CDATA[Six months of public GitHub data across 350+ startup organizations. Which commit patterns actually predict a Series A round? Plus the public Q3 2026 watchlist - bookmark and verify.]]></description>
      <content:encoded><![CDATA[<p>Six months ago I started watching public GitHub for a leading indicator that VCs were missing.</p>
<p>The thesis was simple. Hedge funds spent the last decade extracting alpha from satellite imagery, credit card panels, and shipping data. The venture-capital equivalent - public engineering activity on GitHub - has been sitting in plain sight, ignored by most institutional sourcing teams, who still rely on Crunchbase, warm intros, and Twitter.</p>
<p>I built a crawler. I pointed it at 350+ startup organizations across 15 sectors. I let it run weekly for six months. Here is what I learned.</p>
<h2>What the Signal Actually Looks Like</h2>
<p>The single most predictive indicator is not commit volume. It is commit velocity *change*.</p>
<p>A startup that ships 200 commits per week and continues to ship 200 commits per week tells you nothing. A startup that goes from 80 commits to 240 commits inside 14 days tells you something is happening. Often it is a fundraise close. Often it is a launch. Sometimes it is both.</p>
<p>We track three signals together:</p>
<ol><li>Commit velocity (14-day window) - total commits to the org's most active repo</li><li>Contributor growth (30-day window) - change in the count of unique contributors</li><li>New repo creation (30-day window) - fresh repos appearing in the org</li></ol>
<p>When all three accelerate inside the same two-week window, we classify the startup as &quot;accelerating&quot;. In our backtest across Q3 and Q4 of 2025, roughly 70% of accelerating startups announced a fundraise within six weeks of the signal firing.</p>
<h2>The Four Signal Types</h2>
<p>Not every acceleration looks the same. We classify them into four pre-fundraise patterns:</p>
<p>Engineering hiring burst. Contributor count jumps 40 percent or more inside 30 days. Often pre-Series A - the company has signed term sheets, started hiring, and the new engineers are pushing first commits. We catch this earlier than LinkedIn employee counts because contribution lands before &quot;Senior Engineer at Acme&quot; gets posted.</p>
<p>Infrastructure buildout. Commits to ops, infra, deploy, and observability repos spike. The company is preparing to scale. Usually accompanies a Series A or Series B that is funding go-to-market expansion.</p>
<p>Deploy frequency spike. Commits per day double. Often a sign of a launch run-up. The team is shipping fixes and features fast. Sometimes followed by a product launch and a press cycle, sometimes by a fundraise where the metrics make the deck.</p>
<p>Framework migration. The team is migrating to a new stack - Next.js, Bun, a new ORM, fresh CI. This often happens 60 to 120 days before a Series A. It is the engineering equivalent of cleaning your apartment before parents visit. Our hypothesis: founders know due diligence is coming and want a clean codebase to walk an investor through.</p>
<h2>Where the Signal Fails</h2>
<p>Honesty: the signal is bad for AI-pure startups. They commit constantly regardless of stage. Signal-to-noise is poor. We exclude AI-only orgs from the strongest classification tier and weight other features more heavily.</p>
<p>The signal is also useless for stealth startups that do not open-source. If your dream company is in stealth, GitHub gives you nothing.</p>
<p>And the signal is not investment advice. It tells you who to talk to. It does not tell you who to wire money to. The decision still requires founder conversations, product evaluation, market analysis, and competitive teardown. Engineering velocity is a sourcing signal, not an investment thesis.</p>
<h2>What Does NOT Predict a Round</h2>
<p>Things that look meaningful in the data but turn out to be noise:</p>
<p>- Star count. Vanity. A 30,000-star repo means it had a viral moment. It does not mean revenue.
- Fork count. Lagging indicator at best. By the time forks accumulate, the round is closed.
- Single-repo commit volume. Founders cosplay productivity. One person committing 400 times a week to one repo is a productivity tell, not a fundraise tell.
- The GitHub trending tab. If a project is trending, the round is already 80 percent allocated. You are too late.</p>
<p>The signal is in *change*, not in *level*.</p>
<h2>The Public Watchlist (Q3 2026 Predictions)</h2>
<p>Today I published the <a href="/predicted">10 startups our model predicts will raise in Q3 2026</a>. It is dated. It is bookmarkable. It is falsifiable. Come back in 6 months and check.</p>
<p>Each card on the watchlist links to the underlying GitHub org so you can audit the signal yourself. The free <a href="/predict">/predict tool</a> lets you score any other startup with the same engine in seconds.</p>
<p>I am publishing this in public on purpose. Backtests are easy to fake. Live forward-tested predictions are hard to fake. The only way to build trust in an alternative data source is to show your work and let the future judge it.</p>
<h2>How to Use This in Practice</h2>
<p>If you are a VC, scout, or angel:</p>
<ol><li>Subscribe to the free weekly digest at <a href="https://gitdealflow.com/#signup">gitdealflow.com</a>. Five breakouts every Monday morning, in your inbox before any other source has them.</li><li>Bookmark the <a href="/predicted">Q3 2026 watchlist</a>. When something on the list raises, you will know we called it.</li><li>Use the <a href="/predict">/predict tool</a> before every founder meeting. Two seconds of GitHub-signal context can save a 30-minute call with a company whose engineering momentum is already cooling.</li></ol>
<p>If you are a founder:</p>
<ol><li>Run <a href="/predict">/predict</a> on your own org. Know what your public engineering signal looks like to investors before the next pitch.</li><li>If you are pre-fundraise and your signal is steady or decelerating, fix it before you start the round. Investors who use alt-data are reading this in real time.</li><li>If you are stealth, you are invisible to this signal. Trade-off acknowledged.</li></ol>
<h2>What is Next</h2>
<p>We refresh the dataset every Monday at 09:00 UTC. New watchlists every quarter. Methodology updates published at <a href="/methodology">/methodology</a> when the underlying classification thresholds change.</p>
<p>If you find a startup we should be tracking that is not in the index yet, the <a href="/predict">/predict</a> tool tells you so directly and lets you submit it.</p>
<p>The first VC to systematically use engineering-velocity data as a sourcing input has already won the next decade of seed-stage deal flow. The question is who that is going to be.</p>]]></content:encoded>
      <pubDate>Sun, 19 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[7 Startup Engineering Metrics Every Investor Should Track]]></title>
      <link>https://signals.gitdealflow.com/blog/startup-engineering-metrics-investors-should-track</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/startup-engineering-metrics-investors-should-track</guid>
      <description><![CDATA[Seven engineering metrics from public GitHub data that help investors evaluate startup momentum: commit velocity, contributor growth, repo expansion, weekend activity, and more. A practical checklist for data-driven deal sourcing.]]></description>
      <content:encoded><![CDATA[<p>Most investors evaluate startups on revenue growth, market size, and team pedigree. These are important. But they are also the metrics that every other investor looks at.</p>
<p>Engineering metrics - available from public GitHub data - provide a complementary view of startup health that almost nobody monitors. Here are seven metrics worth tracking, what they tell you, and how to use them. For the broader context on why this data matters, see <a href="/blog/what-is-deal-flow-signal">what is deal flow signal</a>.</p>
<h2>1. What Is Commit Velocity?</h2>
<p>What it is: Total commits to a startup's most active public repository over a rolling 14-day window.</p>
<p>What it tells you: The raw volume of engineering output. High absolute velocity is not inherently meaningful - some teams commit frequently with small changes, others commit less often with larger changes. The value is in tracking it over time to establish a baseline.</p>
<p>How to use it: Commit velocity is the denominator for the most important metric (velocity change). Track it to understand the startup's normal operating rhythm before evaluating whether a change is significant.</p>
<h2>2. Why Is Commit Velocity Change the Primary Signal?</h2>
<p>What it is: The percentage change in commit velocity compared to the preceding 14-day window.</p>
<p>What it tells you: Whether the engineering team is accelerating, maintaining pace, or slowing down. This is the single most useful engineering metric for investors because it measures *acceleration* - the rate of change.</p>
<p>How to use it: A sustained velocity change above +50% for 3+ consecutive windows is a meaningful signal. At VC Deal Flow Signal, this is the primary ranking metric across all 15 sectors. Startups showing +100% or higher velocity change are flagged as top movers.</p>
<p>Benchmark: In our dataset, the average velocity change across 43 tracked startups is approximately +15%. Anything above +50% is unusual. Above +100% is a regime change.</p>
<h2>3. What Does Contributor Count Tell Investors?</h2>
<p>What it is: The number of unique contributors to the startup's most active public repository.</p>
<p>What it tells you: A rough proxy for engineering team size. More useful as a trend than an absolute number, since not all employees contribute to public repos, and not all contributors are employees.</p>
<p>How to use it: Compare contributor count to the startup's claimed team size. A company claiming 30 engineers but showing 5 GitHub contributors either has most code in private repos (normal) or is overstating their team (investigate further).</p>
<h2>4. Why Does Contributor Growth Rate Predict Fundraises?</h2>
<p>What it is: The change in unique contributor count over a 6-week comparison window.</p>
<p>What it tells you: Whether the engineering team is growing. A sudden jump (50%+ in a short window) almost always indicates a hiring burst - new engineers who joined and started committing.</p>
<p>How to use it: Contributor growth above 50% in a 2-week window is our most reliable fundraise predictor. New hires start committing code within days of joining. If you see the contributor count step up sharply, the round likely closed recently and the announcement is coming.</p>
<h2>5. What Does New Repository Creation Signal?</h2>
<p>What it is: The number of public repositories created by the startup's GitHub organization in the last 30 days.</p>
<p>What it tells you: Whether the startup is expanding its technical surface area. New repos usually mean new microservices, SDKs, internal tools, or platform components.</p>
<p>How to use it: Three or more new repos in 30 days is what we call an &quot;infrastructure buildout&quot; signal. This pattern is classic Series A behavior: the core product works, and the team is building the surrounding platform. It signals both technical maturity and available capital.</p>
<h2>6. What Does Weekend Commit Activity Reveal?</h2>
<p>What it is: The proportion of commits that occur on Saturday and Sunday versus weekdays.</p>
<p>What it tells you: How intensely the team is working. Weekend commits from multiple contributors (not just a solo founder) indicate a deadline push.</p>
<p>How to use it: A sustained shift from weekday-only to 7-day commit patterns across multiple contributors is a soft signal that something time-sensitive is happening: a product launch, a fundraise demo, or a competitive response. This metric is most useful as a confirming signal alongside velocity change.</p>
<h2>7. What Do Language and Framework Choices Indicate?</h2>
<p>What it is: The programming languages and frameworks visible in the startup's public repositories.</p>
<p>What it tells you: The technical stack and maturity level. A seed-stage startup using Kubernetes, Terraform, and enterprise monitoring tools may be over-engineering. A growth-stage company still using prototype-quality tools may have hidden technical debt.</p>
<p>How to use it: Cross-reference with the startup's claimed technology during due diligence. If they say they are building an AI platform, their repos should show Python, ML frameworks, and data processing infrastructure. If the repos tell a different story, ask why.</p>
<h2>How Should Investors Use These Metrics Together?</h2>
<p>When evaluating a startup from its public GitHub profile, check these in order:</p>
<ol><li>Is commit velocity change positive and above 50%? (Active acceleration)</li><li>Has contributor count grown recently? (Team scaling)</li><li>Are there new repos in the last 30 days? (Platform building)</li><li>Is the activity product-related, not just docs/CI/CD? (Meaningful work)</li><li>Does the tech stack match the company's pitch? (Consistency check)</li></ol>
<p>If a startup passes all five checks, it is worth a deeper look. If it fails the first two, move on - the engineering signal is not there. For real-world application, read how investors <a href="/blog/github-due-diligence-for-vcs">use GitHub for technical due diligence</a>.</p>
<p>VC Deal Flow Signal automates checks 1-4 across 15 sectors weekly. Browse the sector rankings to see which startups pass the screen right now.</p>]]></content:encoded>
      <pubDate>Tue, 14 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[What Is Engineering Acceleration? The Metric VCs Are Starting to Track]]></title>
      <link>https://signals.gitdealflow.com/blog/what-is-engineering-acceleration</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/what-is-engineering-acceleration</guid>
      <description><![CDATA[Engineering acceleration is the rate of change in a startup's public GitHub commit velocity, contributor count, and repository activity, not participation in an accelerator program like Y Combinator. Learn why this metric matters more than absolute commit counts and how investors use it to time fundraise signals.]]></description>
      <content:encoded><![CDATA[<p>Engineering acceleration is the single most important metric at VC Deal Flow Signal, and the reason the signal works as a deal sourcing tool. This page is the definition piece. For the full operating playbook covering pipeline, benchmarks, and predictive analytics, see <a href="/blog/how-vcs-track-engineering-acceleration-2026-playbook">How VCs Track Startup Engineering Acceleration: The Complete 2026 Playbook</a>.</p>
<p>A vocabulary note up front. Throughout this site, engineering acceleration always refers to code-side momentum measured from public GitHub activity, commit velocity, contributor growth, repository creation. It has nothing to do with startup accelerator programs like Y Combinator or Techstars, despite the unfortunate vocabulary collision. When the word &quot;accelerator&quot; appears here without qualification, it is shorthand for the GitHub-derived signal, not for a program.</p>
<h2>What is engineering acceleration?</h2>
<p>Engineering acceleration measures the rate of change in a startup's engineering output. Not how much code they write, but how much faster they are writing it compared to their own baseline.</p>
<p>The formula is straightforward: take the 14-day commit count for a startup's most active public repositories, compare it to the prior 14-day window, and express the change as a percentage. A startup with 40 commits this period and 20 last period shows +100% acceleration. A startup with 200 commits this period and 220 last period shows -10%, even though the absolute volume is higher than the +100% case.</p>
<p>This is different from absolute engineering volume. A company with 500 commits per week is not necessarily more interesting than one with 50, what matters is whether the 50-commit company just jumped from 25. That jump is the signal. Absolute volume reflects team size, codebase maturity, and shipping conventions, all of which differ across companies in ways that obscure comparison. Rate of change normalizes against a startup's own historical baseline, which is the most honest comparable.</p>
<p>The metric extends beyond commit velocity. Three other dimensions carry independent information: contributor count change (new engineers being added), repository expansion (new product surfaces), and language mix shift (architectural transitions). Together the four metrics describe an acceleration pattern that can be classified into operational types, see the four signal types section below.</p>
<h2>Why does this metric matter for investors?</h2>
<p>The causal chain that makes engineering acceleration useful for investors is short. A startup decides to raise capital, or has just closed a round, or plans a major launch. That decision drives engineering activity: hiring engineers, sprinting toward a milestone, building new infrastructure. The engineering activity produces commits, pull requests, and new repositories. The activity is observable in GitHub's public data within hours or days of happening. Press coverage, Crunchbase entries, SEC Form D filings, and LinkedIn announcements follow weeks to months later.</p>
<p>Most investors only see the downstream signals, the press release, the database entry, the hiring announcement. By the time a startup appears in those sources, the round is often allocated and the deal is competitive. Engineering acceleration shows up before any of that. The lead time across the 350+-startup panel maintained at VC Deal Flow Signal is a median of three to six weeks before a public fundraise announcement <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>.</p>
<p>The economic case for investors is compounding. Warm outreach during the pre-fundraise window converts at meaningfully higher rates than cold outreach during a competitive round. Even at the angel and seed level, where check sizes are small, being the first thoughtful conversation a founder has had about a round is worth disproportionately more than being the tenth. The metric is, in practice, a top-of-funnel sourcing tool that turns the global startup population into a tractable weekly screen.</p>
<p>The metric is not magic. Many accelerations resolve in disappointing ways: the team ships a launch and stalls, the round is extended rather than upsized, the apparent burst was a single contributor's hackathon. The framework's job is not to eliminate false positives, it is to make them tractable. The output is a screen, not a buy decision.</p>
<h2>How is engineering acceleration different from DORA metrics?</h2>
<p>DORA metrics, deployment frequency, lead time for changes, change failure rate, and time to restore, are the four canonical measures of engineering process quality popularized by Google's DORA research <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>. They measure how reliably and quickly a team ships, with an emphasis on production safety and recovery. DORA is the standard internal toolkit for engineering managers and CTOs.</p>
<p>Engineering acceleration measures something different: the rate of change in engineering output volume. Not how reliably the team ships, but whether the team is speeding up. The two metrics answer different questions. DORA answers &quot;is this engineering organization healthy?&quot; Acceleration answers &quot;is something meaningful changing in this team's pace?&quot;</p>
<p>The practical difference matters for who can use each metric. DORA requires internal access to CI/CD pipelines, deployment systems, and incident tooling. It is unobservable from outside the company. Engineering acceleration is computed entirely from public GitHub data, commits, contributors, repositories, which means anyone with API access can compute it for any public organization. That asymmetry is what makes acceleration useful as an investor signal: it can be applied at scale to companies the investor has no relationship with.</p>
<p>Both metrics are valuable in their domains. A founder optimizing for DORA scores is improving engineering process. An investor watching for acceleration is identifying changes in startup momentum. They sometimes correlate, a team scaling up will often improve DORA metrics and show acceleration simultaneously, but the metrics are not substitutes.</p>
<h2>What are the four signal types?</h2>
<p>When a startup shows acceleration, the pattern can be classified into one of four operational types based on which underlying metrics are moving. The classification matters because each pattern implies a different diligence question.</p>
<p>The hiring burst is the pattern most strongly correlated with a recent or imminent fundraise. The fingerprint is rising commit velocity combined with rising unique contributor count, both moving in the same direction at meaningful magnitude. The detection rule used in the public methodology <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup> is contributor count up at least 30 percent and commit velocity up at least 60 percent in the same 14-day window, sustained into a second period. Hiring bursts almost always reflect committed capital, because adding engineers requires payroll commitments, which require runway visibility, which usually requires recent or imminent fundraise activity.</p>
<p>The shipping sprint is velocity rising while contributor count stays flat. This pattern signals a launch push, the existing team is pushing harder toward a milestone. Detection rule: velocity up at least 100 percent with contributor count change under 15 percent. Sprints are interesting investor signals but require different conversations: the team is preparing for something, often a launch, sometimes a fundraise narrative built around the launch. Catching a sprint early gives investors a chance to engage with the founder before the launch event.</p>
<p>The infrastructure buildout is repository creation accelerating relative to historical baseline. This signals architectural investment, platform migrations, new product surfaces, or build-out of internal tooling. Detection rule: at least three new repositories created in 30 days versus a prior 30-day baseline of zero. Buildouts often presage Series A or Series B fundraises because they imply the team is committing to scale-stage investment in technical foundations.</p>
<p>The platform migration is language mix shifting between primary languages over a quarter. This is the slowest-moving but most strategically significant signal, it implies the team is committing to a new technical direction. Detection rule: at least 20 percentage points of language mix migrating between primary languages over a 90-day window. Migrations often coincide with senior engineering hires whose stack preferences drive architectural choices, or with platform rebuilds that the team has been planning for months.</p>
<p>Each signal type has implications for how an investor should approach the founder. A hiring burst suggests asking about recent or imminent capital. A shipping sprint suggests asking about the upcoming launch and its dependencies. An infrastructure buildout suggests asking about architectural strategy. A platform migration suggests asking about the technical bet driving the change. The signal types direct the investor's attention; the actual investment decision still requires founder conversations.</p>
<p>For full definitions of each signal type and the detection rules, see the <a href="/glossary">glossary</a>. For weekly rankings of startups currently showing each pattern, see the <a href="/">sector rankings</a>.</p>
<h2>How is acceleration measured in practice?</h2>
<p>The measurement pipeline at VC Deal Flow Signal has four stages: ingestion, normalization, aggregation, and rate-of-change computation.</p>
<p>Ingestion pulls weekly data from the GitHub REST API <sup><a href="https://signals.gitdealflow.com/blog/#references">[2]</a></sup> for approximately 350+ startup organizations across 15 sectors. The relevant endpoints are repositories list, commit activity, contributors, and releases. Free authenticated rate limits (5,000 requests per hour per token) are sufficient for the panel size when paced over six hours.</p>
<p>Normalization removes the most common noise sources. Bot accounts (Dependabot, Renovate, GitHub Actions) can drive double-digit commit counts per week without any human engineering activity, so commits authored by accounts matching common bot patterns are excluded. A second normalization layer filters commits by file count and diff size, removing trivial commits that inflate counts without reflecting meaningful engineering work.</p>
<p>Aggregation produces four time series per organization: commit velocity (count over rolling 14 days), unique contributor count over the same window, repositories created in the period, and language mix as a percentage breakdown of commit volume by primary language. Each is computed weekly and stored with timestamps.</p>
<p>Rate-of-change computation produces the headline acceleration number plus the four pattern-classification metrics. The acceleration number is the percentage change in 14-day commit velocity versus the prior 14-day window. The pattern classification compares the four core metrics against the detection rules above.</p>
<p>The full methodology is published openly. The peer-style write-up is on SSRN <sup><a href="https://signals.gitdealflow.com/blog/#references">[3]</a></sup>; the working dataset is mirrored on Zenodo and OpenAlex; the production pipeline is documented in the <a href="/methodology">methodology</a> page on this site. Reproducibility is intentional, investors should be able to evaluate the signal quality before acting on it.</p>
<h2>Common pitfalls in interpreting acceleration</h2>
<p>The framework's failure modes are predictable, and a working pipeline is one whose users have internalized them.</p>
<p>The bot inflation problem affects organizations with aggressive automation tooling. The fix is bot exclusion, applied consistently. The single-contributor problem affects pre-seed startups, where a solo founder can drive +500% acceleration through pure personal effort. The two-period confirmation rule plus contributor cross-validation handles most cases, but pre-seed signals always require human review.</p>
<p>The launch-versus-fundraise problem is a recurring interpretation issue. Both events are interesting, but they imply different conversations. The diagnostic clue is sequencing: a launch tends to produce a coordinated burst over four to six weeks followed by a sharp drop-off, while a fundraise-driven acceleration sustains over a longer period. The acquisition-driven acceleration affects later-stage companies, where absorbing an acquired engineering team produces +200% acceleration unrelated to internal momentum.</p>
<p>For a comprehensive treatment of pitfalls and the diligence checks that catch them, see the <a href="/blog/how-vcs-track-engineering-acceleration-2026-playbook">common pitfalls section in the cornerstone playbook</a>.</p>
<h2>How investors use the metric</h2>
<p>The pragmatic use of engineering acceleration is a weekly digest. Every Monday, an investor reviews the top-ranked breakouts in their sectors of interest, cross-references against their CRM for any prior contact, prioritizes unflagged ones for a 30-minute desk dive, and tags the rest for monitoring. Total time investment is roughly 30 minutes per week, which is small enough that the metric is sustainable as part of a normal sourcing workflow.</p>
<p>The signal complements rather than replaces existing sourcing channels. Founder networks, demo days, and accelerator pipelines stay intact; engineering acceleration adds an external, quantitative top-of-funnel feed that is hard to source any other way. The most successful fund deployments treat the signal as a conversation prompt, not a buying decision: the first outreach acknowledges the team's pace, asks what is driving it, and is open to whatever the founder wants to share.</p>
<p>For investors getting started, the easiest entry point is the <a href="https://gitdealflow.com">free weekly Signal Report</a>, which surfaces the top five breakout startups across all sectors every Monday. The Dashboard at EUR 49/month adds full sector, stage, and geography filters. The MCP server (npx -y @gitdealflow/mcp-signal) provides programmatic access for funds with engineering capacity. The full operating playbook is at <a href="/blog/how-vcs-track-engineering-acceleration-2026-playbook">How VCs Track Startup Engineering Acceleration: The Complete 2026 Playbook</a>.</p>
<p>Engineering acceleration is not a magic source of alpha. It is a well-defined, publicly observable signal that arrives weeks before the data sources most investors currently use. Capturing the timing edge requires nothing more than the discipline to look at the signal every week.</p>
<p>Browse the <a href="/">sector rankings</a> to see which startups are showing engineering acceleration on GitHub right now.</p>]]></content:encoded>
      <pubDate>Tue, 14 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Commit Velocity Explained: What Investors Need to Know]]></title>
      <link>https://signals.gitdealflow.com/blog/commit-velocity-explained</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/commit-velocity-explained</guid>
      <description><![CDATA[Commit velocity is the total number of commits to a startup's GitHub repository over a rolling 14-day window. Learn what it measures, what it misses, and how to interpret it for deal sourcing.]]></description>
      <content:encoded><![CDATA[<p>Commit velocity is one of the most cited - and most misunderstood - metrics in GitHub-based deal sourcing.</p>
<h2>What Is Commit Velocity?</h2>
<p>Commit velocity is the total number of commits to a startup's most active public GitHub repository over a rolling 14-day window. At VC Deal Flow Signal, we pull this data from the GitHub API's commit activity endpoint <sup><a href="https://signals.gitdealflow.com/blog/#references">[1]</a></sup>.</p>
<p>The metric is intentionally simple: count the commits. No weighting by lines of code, no filtering by author, no adjustment for commit size. Simplicity makes it comparable across companies and sectors.</p>
<h2>What Does Commit Velocity Actually Measure?</h2>
<p>Commit velocity measures engineering output volume - how much work is being pushed to version control. It is a proxy for engineering activity, not engineering quality.</p>
<p>A startup with 200 commits in 14 days has roughly 14 commits per day. For a team of 10 engineers, that is a healthy shipping cadence. For a solo founder, it might indicate automated tooling or excessive granularity.</p>
<h2>What Are the Limitations?</h2>
<p>Commit velocity has known limitations that investors should understand:</p>
<p>Commit size varies: One commit might change a single config line; another might refactor 5,000 lines. Velocity treats them equally.</p>
<p>Automation inflates counts: CI/CD bots, automated dependency updates, and auto-formatting tools can inflate commit counts without meaningful engineering work.</p>
<p>Private repos are invisible: Many startups keep their core product code in private repositories. Commit velocity only captures public activity.</p>
<p>Squash vs. merge: Teams that squash commits will show lower velocity than teams that merge individual commits. This is a workflow choice, not a quality signal.</p>
<h2>Why Commit Velocity Change Matters More</h2>
<p>This is the key insight: absolute commit velocity is noisy. Commit velocity change - the rate at which velocity is accelerating - is the real signal. See our detailed guide on <a href="/blog/how-to-read-github-signals-for-startup-investing">how to read GitHub signals for startup investing</a>.</p>
<p>A startup going from 20 to 40 commits in a 14-day window shows +100% velocity change. That acceleration has meaning regardless of the absolute numbers - something changed in how the team is working. That something is what investors care about.</p>
<p>Visit our <a href="/glossary#commit-velocity">glossary</a> for formal definitions of all engineering metrics.</p>]]></content:encoded>
      <pubDate>Mon, 13 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Pre-Seed Deal Sourcing with GitHub Data: A Practical Guide]]></title>
      <link>https://signals.gitdealflow.com/blog/pre-seed-deal-sourcing-github</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/pre-seed-deal-sourcing-github</guid>
      <description><![CDATA[How to use GitHub engineering signals to find pre-seed startups before they raise. Covers what pre-seed activity looks like on GitHub, signal patterns, and a step-by-step sourcing workflow.]]></description>
      <content:encoded><![CDATA[<p>Pre-seed startups are invisible to most deal sourcing tools. They have no Crunchbase entry, no press coverage, no PitchBook profile. But many of them have GitHub activity.</p>
<h2>Why GitHub Works for Pre-Seed Sourcing</h2>
<p>GitHub is one of the few public data sources where pre-seed activity is visible. Before a startup has a pitch deck, before it has a website, the founders are writing code. That code - if the repositories are public - creates a data trail.</p>
<p>At VC Deal Flow Signal, we estimate startup stage from contributor count: pre-seed companies typically have 1-7 contributors. When we see a small team showing disproportionate acceleration - committing at 3-5x their baseline rate - something is happening worth investigating.</p>
<h2>What Pre-Seed Signals Look Like</h2>
<p>Pre-seed engineering activity has a distinctive pattern:</p>
<p>Low absolute velocity, high acceleration: A solo founder going from 5 commits/week to 25 commits/week shows +400% velocity change. In absolute terms, 25 commits is nothing. But the acceleration is the signal.</p>
<p>Infrastructure buildout from zero: New repositories appearing (3+ in 30 days) in a young organization suggests the founder is moving from prototype to structured development. This is classic pre-seed-to-seed transition behavior.</p>
<p>Contributor count jumping from 1 to 3-4: When a solo founder suddenly has co-contributors, they either found a co-founder, hired their first engineer, or attracted open source contributors. All three are positive signals.</p>
<h2>A Pre-Seed Sourcing Workflow</h2>
<ol><li>Check the <a href="/">sector rankings</a> weekly and filter for startups showing &quot;Pre-seed&quot; stage estimation</li><li>Look for companies with +200% or higher velocity change from a small base</li><li>Open their GitHub organization - is the activity product-related?</li><li>Check if the founder is active on Twitter, Hacker News, or Indie Hackers</li><li>If signals align, reach out before anyone else knows the company exists</li></ol>
<p>For the complete screening methodology, see the <a href="/blog/startup-engineering-metrics-investors-should-track">7 engineering metrics every investor should track</a>.</p>]]></content:encoded>
      <pubDate>Sun, 12 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Series A Signals: What GitHub Data Reveals About Growth-Stage Startups]]></title>
      <link>https://signals.gitdealflow.com/blog/series-a-signals-github-data</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/series-a-signals-github-data</guid>
      <description><![CDATA[Series A startups show distinctive GitHub patterns: infrastructure buildout, rapid contributor growth, and platform expansion. Learn what these signals mean for investors evaluating growth-stage deals.]]></description>
      <content:encoded><![CDATA[<p>Series A startups look different on GitHub than pre-seed or seed companies. The patterns are distinctive enough to identify from public data alone.</p>
<h2>What Makes Series A GitHub Activity Different</h2>
<p>At the pre-seed and seed stages, GitHub activity is concentrated: one or two repositories, a small team, and commit patterns driven by individual contributors. At Series A, the picture changes.</p>
<p>The defining characteristic is platform expansion. The core product works - customers are using it - and now the team is building everything around it: SDKs, developer documentation, internal tools, deployment infrastructure, monitoring systems.</p>
<h2>The Infrastructure Buildout Signal</h2>
<p>The strongest Series A signal is infrastructure buildout: 3 or more new public repositories created in 30 days. This is not a founder experimenting with side projects. This is a company with capital deploying it into platform development.</p>
<p>Common new repositories at this stage include API client libraries, CLI tools, integration frameworks, and documentation sites. Each represents a deliberate investment in making the product accessible to more users or developers.</p>
<h2>Contributor Growth as a Post-Raise Indicator</h2>
<p>When contributor count jumps 50% or more in a short window, the company has likely just closed a round and is scaling the engineering team. At Series A, this typically means going from 8-12 contributors to 15-25.</p>
<p>The timing is important: contributor growth appears in GitHub data within weeks of new engineers joining, but the fundraise announcement may not appear on Crunchbase for another 6-12 weeks. This gap is the investor's opportunity.</p>
<h2>How to Use These Signals</h2>
<p>Filter the <a href="/">sector rankings</a> for startups estimated at &quot;Series A/B&quot; stage. Look for companies showing &quot;Infrastructure buildout&quot; signal type with contributor growth above 50%. Cross-reference with the <a href="/trending">trending page</a> to find the strongest movers.</p>
<p>For a complete framework on interpreting these signals, see our guide to <a href="/blog/github-due-diligence-for-vcs">GitHub due diligence for VCs</a>.</p>]]></content:encoded>
      <pubDate>Sat, 11 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[How to Source Startup Deals Before They Appear on Crunchbase]]></title>
      <link>https://signals.gitdealflow.com/blog/source-startup-deals-before-crunchbase</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/source-startup-deals-before-crunchbase</guid>
      <description><![CDATA[Crunchbase tells you what already happened. Learn three approaches to finding startups before they raise - using GitHub signals, community sourcing, and hiring data as leading indicators.]]></description>
      <content:encoded><![CDATA[<p>Every investor uses Crunchbase. That is exactly the problem.</p>
<p>Crunchbase is excellent at what it does: a comprehensive database of startup funding rounds, team members, and company profiles. But by design, it is a lagging indicator. A company appears in your Crunchbase alert after the round closes, after the terms are set, after the press release is written. You are seeing what already happened.</p>
<p>The investors who consistently get into the best deals are the ones who found the company before it appeared on Crunchbase. This post covers three practical approaches to doing that.</p>
<h2>How Do GitHub Engineering Signals Help You Find Deals First?</h2>
<p>GitHub engineering activity is the earliest publicly available signal of startup momentum - and part of a broader shift toward <a href="/blog/alternative-data-venture-capital">alternative data in venture capital</a>. The logic is straightforward: engineering acceleration precedes product milestones, which precede fundraise decisions, which precede Crunchbase entries.</p>
<p>When a startup's commit velocity doubles in a two-week window and the change is sustained, something fundamental has shifted. Common causes:</p>
<p>- Post-fundraise scaling: New capital deployed → new engineers hired → commit velocity spikes. The round closed but is not yet announced. Lead time: 6-12 weeks before the Crunchbase entry.
- Product-market fit iteration: Customer feedback driving rapid feature development. Lead time: 8-16 weeks before a fundraise decision is even made.
- Launch preparation: Team pushing toward a release. Often followed by press coverage and investor attention.</p>
<p>What to look for:
- Commit velocity change &gt; 100%: The startup's 14-day commit count doubled compared to the prior window.
- Contributor growth &gt; 50%: New team members appeared - likely recent hires.
- 3+ new repositories in 30 days: Infrastructure buildout, classic Series A behavior.</p>
<p>This is what VC Deal Flow Signal tracks across 15 sectors weekly. The top movers consistently include companies that announce raises 4-8 weeks later.</p>
<h2>Which Community Platforms Surface Startups Earliest?</h2>
<p>Community platforms surface startups at different stages of visibility:</p>
<p>Hacker News Show HN - Very early signal. Founders posting technical projects before they have a pitch deck. Lead time: months before any institutional awareness. The challenge is volume - most Show HN posts are weekend projects, not fundable companies.</p>
<p>Indie Hackers - Build-in-public culture means founders share revenue numbers, growth metrics, and technical decisions openly. Lead time: weeks to months. The signal is in the engagement - posts that generate deep technical discussion often indicate real traction.</p>
<p>Product Hunt - Launch signal, not traction signal. By the time a startup launches on Product Hunt, they usually have a polished product and some early customers. Lead time: 2-4 weeks before broader awareness.</p>
<p>Y Combinator batch lists - Published at demo day, which is late in the cycle (investors already competing for these companies). But the companies that raise quietly before or after demo day are the ones to watch.</p>
<p>The community sourcing approach works best when you are deeply embedded in a specific community. An investor who reads r/venturecapital daily catches signals that a broader scan would miss.</p>
<h2>How Can Hiring Data Reveal Upcoming Fundraises?</h2>
<p>Job postings reveal a startup's growth plans before they are announced publicly:</p>
<p>- Senior engineering hires (VP Engineering, Staff Engineer): Team is scaling, likely post-fundraise.
- Head of Sales / VP Marketing: Go-to-market is being built. Product-market fit is likely established.
- Multiple simultaneous postings: Coordinated hiring push, usually funded by a recent or imminent round.</p>
<p>Where to find hiring signals:
- LinkedIn job postings (filter by company size 1-50)
- AngelList/Wellfound job boards
- Y Combinator's Work at a Startup
- Hacker News monthly &quot;Who's Hiring&quot; threads</p>
<p>Lead time: 4-8 weeks before the round is announced. Shorter than GitHub signals, but the signal is more explicit about the type of growth.</p>
<h2>How Should Investors Combine All Three Signal Types?</h2>
<p>The most effective approach combines all three signal types:</p>
<ol><li>GitHub signals surface companies showing engineering acceleration (earliest warning)</li><li>Community signals add context - is the founder talking about traction? Customer feedback? Hiring?</li><li>Hiring signals confirm the growth trajectory - are they actively building the team?</li><li>Crunchbase verifies funding history and competitive landscape (due diligence, not sourcing)</li></ol>
<p>This progression gives you the best of both worlds: timing advantage from alternative data, and verification depth from traditional sources.</p>
<h2>What Does This Look Like in Practice?</h2>
<p>Every week, check the sector rankings for your focus areas. When an unfamiliar name appears in the top 3 with a strong acceleration signal:</p>
<ol><li>Spend 5 minutes on their GitHub - is the activity product-related or maintenance noise?</li><li>Search Hacker News, Reddit, and Twitter for the company name - any community buzz?</li><li>Check their careers page - are they hiring?</li><li>Open Crunchbase - what is their funding history? Are they pre-raise?</li><li>If all signals align, reach out to the founder.</li></ol>
<p>This workflow takes 15-20 minutes per company and puts you weeks ahead of investors who only use Crunchbase alerts. For the full screening checklist, see the <a href="/blog/startup-engineering-metrics-investors-should-track">7 engineering metrics every investor should track</a>.</p>
<p>Browse the sector rankings to start identifying startups before they appear in your inbox.</p>]]></content:encoded>
      <pubDate>Fri, 10 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Open Source Startups: An Investor's Guide to GitHub Signal Analysis]]></title>
      <link>https://signals.gitdealflow.com/blog/open-source-startups-investor-guide</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/open-source-startups-investor-guide</guid>
      <description><![CDATA[Open source startups present unique challenges for GitHub-based deal sourcing. Learn how to separate community contributions from commercial engineering activity and identify the open source companies worth investing in.]]></description>
      <content:encoded><![CDATA[<p>Open source startups are some of the most interesting investment opportunities in developer tools and infrastructure. But they present unique challenges for GitHub-based signal analysis.</p>
<h2>The Community Noise Problem</h2>
<p>For most startups, commit velocity is a clean signal - the commits come from the team. For open source startups, commits come from everywhere: core team, community contributors, one-time bug fixers, documentation translators, and bots.</p>
<p>This noise inflates the standard metrics. A popular open source project might show 500 contributors, but only 10 are paid employees. A commit velocity spike might reflect a documentation sprint by the community, not product acceleration by the company.</p>
<h2>How to Separate Commercial from Community Signals</h2>
<p>The key is to focus on the company-owned organization, not the project repository. Most commercial open source startups have a GitHub organization with multiple repos: the main project, plus commercial tools, SDKs, enterprise features, and infrastructure.</p>
<p>Track these separately:
- Core project repo: Community engagement signal (stars, forks, external PRs)
- Organization-level repos: Commercial engineering signal (new repos, internal tools, enterprise features)
- Contributor growth in core maintainers: Hiring signal (new team members with commit access)</p>
<h2>The Strongest Open Source Investment Signal</h2>
<p>The most compelling signal is simultaneous community growth and commercial acceleration. When the open source project is gaining stars and contributors while the company organization is building enterprise infrastructure, the flywheel is working.</p>
<h2>Practical Screening</h2>
<p>Browse the <a href="/startups-to-watch/developer-tools-q2-2026">Developer Tools sector rankings</a> and look for companies with high contributor counts (50+) but moderate contributor growth. Then check if they show infrastructure buildout signals - new repos for commercial features.</p>
<p>For the full deal sourcing framework, see <a href="/blog/source-startup-deals-before-crunchbase">how to source startup deals before they appear on Crunchbase</a>.</p>]]></content:encoded>
      <pubDate>Thu, 09 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[GitHub Signals vs Hiring Data: Which Predicts Fundraises Better?]]></title>
      <link>https://signals.gitdealflow.com/blog/github-signals-vs-hiring-data</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/github-signals-vs-hiring-data</guid>
      <description><![CDATA[Compare GitHub engineering signals and hiring data as leading indicators of startup fundraises. Lead time, reliability, coverage, and which investors should use - or whether the combination beats either alone.]]></description>
      <content:encoded><![CDATA[<p>Two alternative data sources dominate the conversation about startup deal sourcing: GitHub engineering signals and hiring data. Both claim to predict fundraises before traditional channels. Which actually works better?</p>
<h2>GitHub Signals: The Earliest Public Indicator</h2>
<p>GitHub engineering acceleration appears 6-12 weeks before fundraise announcements. The logic: a startup accelerates engineering output → builds product → raises capital → announces the round. GitHub catches step one.</p>
<p>The signal is in the rate of change. A commit velocity increase of +100% or more, sustained over multiple weeks, indicates something fundamental shifted in how the team is working. See <a href="/blog/what-is-engineering-acceleration">what is engineering acceleration</a> for the full explanation.</p>
<h2>Hiring Data: The Most Explicit Indicator</h2>
<p>Job postings on LinkedIn, AngelList, and company career pages provide 4-8 weeks of lead time. Hiring data is later than GitHub signals but more explicit: a &quot;VP Engineering&quot; posting tells you they are scaling the technical team, while a &quot;Head of Sales&quot; posting tells you they are building go-to-market.</p>
<p>Hiring data answers &quot;what are they building?&quot; GitHub data answers &quot;how fast are they building?&quot;</p>
<h2>The Comparison</h2>
<p>Lead time: GitHub wins (6-12 weeks vs 4-8 weeks). Engineering acceleration precedes hiring decisions because teams ship faster before they staff up.</p>
<p>Signal explicitness: Hiring wins. A job posting for &quot;Senior ML Engineer&quot; tells you more about strategic direction than a commit velocity spike.</p>
<p>Coverage: Hiring wins for breadth (every company hires). GitHub wins for depth (commit-level granularity on technical startups).</p>
<p>Cost: Both are free for basic analysis. GitHub data is available via API; hiring data requires scraping or paid platforms.</p>
<h2>The Optimal Approach</h2>
<p>Use both, sequentially. GitHub signals surface the candidates (earliest warning). Hiring data confirms the trajectory (what type of growth). This is the workflow described in our guide on <a href="/blog/source-startup-deals-before-crunchbase">sourcing deals before Crunchbase</a>.</p>
<p>Browse the <a href="/trending">trending page</a> for the startups showing the strongest GitHub engineering acceleration (commit velocity + contributor growth) this week.</p>]]></content:encoded>
      <pubDate>Wed, 08 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Alternative Data for Venture Capital: Why GitHub Is the Most Underused Signal]]></title>
      <link>https://signals.gitdealflow.com/blog/alternative-data-venture-capital</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/alternative-data-venture-capital</guid>
      <description><![CDATA[Alternative data has transformed public market investing. Now it is coming to venture capital. GitHub engineering activity is the most accessible, real-time, and underused alternative data source for startup investors.]]></description>
      <content:encoded><![CDATA[<p>Alternative data changed public market investing over the past decade. Satellite imagery of parking lots, credit card transaction data, app download metrics - hedge funds built entire strategies on signals that traditional analysts ignored.</p>
<p>Venture capital has been slower to adopt alternative data. Most deal sourcing still relies on warm introductions, demo days, and newsletters. The irony is that the most accessible alternative data source for startup investing has been sitting in the open for years: GitHub.</p>
<h2>What Counts as Alternative Data for Venture Capital?</h2>
<p>In public markets, alternative data is any dataset that provides insight into a company's performance beyond traditional financial filings. For venture capital, the concept is similar - any signal that reveals startup traction before it appears through conventional deal sourcing channels.</p>
<p>The main categories of alternative data for VCs:</p>
<p>Engineering activity (GitHub): Commit velocity, contributor growth, repository expansion. Available via public API, updated daily, hard to fake. Lead time: 6-12 weeks before fundraise announcements.</p>
<p>Hiring signals (job boards, LinkedIn): New job postings, especially for senior engineering and go-to-market roles. Scraping required, updated weekly. Lead time: 4-8 weeks.</p>
<p>Web traffic (SimilarWeb, Sensor Tower): Rapid growth in a startup's web or app traffic. Requires paid tools, updated monthly. Lead time: 4-6 weeks.</p>
<p>Social signals (Twitter, HN, Reddit): Mentions, upvotes, and community engagement. Free but noisy, real-time. Lead time: 1-2 weeks (often lagging, not leading).</p>
<p>Patent filings (USPTO, EPO): New patent applications signal R&amp;D direction. Free but delayed by 18 months, so more useful for competitive analysis than timing.</p>
<h2>Why Is GitHub the Best Alternative Data Source for VCs?</h2>
<p>Among all alternative data sources for VCs, GitHub engineering activity has unique properties:</p>
<p>It is continuous and granular. Unlike hiring signals (which appear when a job is posted) or web traffic (which updates monthly), GitHub commits happen daily. You can track weekly velocity changes and catch acceleration patterns in real time.</p>
<p>It is free and public. GitHub's API provides commit history, contributor data, and repository metadata at no cost. No scraping required. No third-party tools needed for basic analysis.</p>
<p>It reflects real work. Commits represent actual engineering output. You cannot game commit velocity the way you can game social media metrics or app store rankings. A team that ships 200 commits in a week did real engineering work.</p>
<p>It reveals intent. The type of engineering activity - new infrastructure repos, contributor scaling, velocity spikes - tells you what phase a startup is in. Infrastructure buildout looks different from feature shipping, which looks different from a documentation sprint before a fundraise.</p>
<h2>How Do Quantitative Investors Approach Alternative Data?</h2>
<p>Quantitative investment firms have understood for years that public data, processed systematically, creates information asymmetry. The edge is not in having exclusive data - it is in reading what others ignore, faster and more consistently.</p>
<p>The same principle applies to venture capital. Every investor has access to GitHub. Almost none of them monitor it systematically. The investor who builds a workflow around engineering signals has a structural timing advantage: they see acceleration patterns 6-12 weeks before the fundraise announcement that fills everyone else's inbox.</p>
<p>This is not theoretical. At VC Deal Flow Signal, we track thousands of startup GitHub orgs across 15 sectors and rank them by engineering acceleration. The patterns are consistent: commit velocity spikes, contributor growth bursts, and infrastructure buildouts appear weeks before TechCrunch writes about the company. We break down the <a href="/blog/5-github-patterns-that-predict-fundraises">5 GitHub patterns that predict fundraises</a> in a separate deep dive.</p>
<h2>How Can Investors Start Using Alternative Data?</h2>
<p>If you are an investor interested in adding alternative data to your sourcing process, start with the highest signal-to-noise ratio source: GitHub engineering acceleration.</p>
<ol><li>Pick 2-3 sectors you know well</li><li>Watch the weekly sector rankings for unfamiliar names in the top 3</li><li>Cross-reference with Crunchbase for funding history and stage</li><li>Reach out to founders during the acceleration window (weeks 2-4 of a velocity spike)</li></ol>
<p>The combination of engineering signals for timing and traditional data for due diligence gives you both a lead time advantage and a solid evaluation framework. For a practical walkthrough, see <a href="/blog/source-startup-deals-before-crunchbase">how to source deals before Crunchbase</a>.</p>
<p>Browse our sector rankings to see which startups are showing engineering acceleration right now, or get the free Signal Report for a weekly summary of the top breakout signals.</p>]]></content:encoded>
      <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Fintech Startup Engineering Signals: What the GitHub Data Shows]]></title>
      <link>https://signals.gitdealflow.com/blog/fintech-startup-engineering-signals</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/fintech-startup-engineering-signals</guid>
      <description><![CDATA[An analysis of engineering acceleration patterns specific to fintech startups. Regulatory-driven development cycles, compliance infrastructure, and what makes fintech GitHub signals different from other sectors.]]></description>
      <content:encoded><![CDATA[<p>Fintech startups show different engineering patterns on GitHub compared to other sectors. Understanding these patterns is essential for investors using engineering signals to source fintech deals.</p>
<h2>Regulatory-Driven Development Cycles</h2>
<p>The most important difference: fintech development cycles are partly driven by regulatory requirements, not just product-market fit. When a fintech startup shows a commit velocity spike, it may be responding to a compliance deadline rather than customer demand.</p>
<p>This does not make the signal less valuable - regulatory compliance requires engineering investment, which requires capital. But it changes the interpretation. A fintech company building KYC infrastructure is not necessarily iterating on product-market fit; it may be preparing for a regulated launch.</p>
<h2>What Infrastructure Buildout Means in Fintech</h2>
<p>In most sectors, infrastructure buildout (3+ new repositories in 30 days) indicates platform expansion. In fintech, the new repositories often serve a different purpose: compliance infrastructure, audit logging, encryption libraries, and regulatory reporting tools.</p>
<p>Look at the repository names and descriptions. New repos named &quot;kyc-service,&quot; &quot;audit-log,&quot; or &quot;compliance-api&quot; tell a different story than &quot;marketplace-sdk&quot; or &quot;developer-tools.&quot;</p>
<h2>The Strongest Fintech Signal</h2>
<p>The most compelling fintech investment signal is simultaneous product acceleration and compliance buildout. When a company is shipping product features and building compliance infrastructure at the same time, it is preparing for a regulated launch. This typically requires significant capital, which means fundraising is imminent or recently completed.</p>
<p>Browse the <a href="/startups-to-watch/fintech-q2-2026">Fintech sector rankings</a> to see the current data, or compare approaches using our <a href="/compare/best-deal-flow-tools-seed-investors">best deal flow tools for seed-stage investors</a> guide.</p>]]></content:encoded>
      <pubDate>Tue, 07 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[AI Startup Engineering Signals in 2026: What Investors Should Watch]]></title>
      <link>https://signals.gitdealflow.com/blog/ai-startup-signals-2026</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/ai-startup-signals-2026</guid>
      <description><![CDATA[The AI sector shows the highest commit velocity of any sector we track. Learn which AI engineering patterns signal real traction vs. hype, and how to use GitHub data to find the AI startups worth investing in.]]></description>
      <content:encoded><![CDATA[<p>AI is the highest-velocity sector in our dataset. It is also the noisiest. Here is how to read AI startup engineering signals in 2026.</p>
<h2>The AI Velocity Paradox</h2>
<p>AI startups show the highest average commit velocity of any sector we track at VC Deal Flow Signal. But high velocity alone does not mean high quality deal flow. The AI sector has more open source experimentation, more research-oriented commits, and more hype-driven activity than any other sector.</p>
<p>The challenge for investors: separating genuine product engineering from research exploration and open source community activity. See our guide on <a href="/blog/open-source-startups-investor-guide">evaluating open source startups</a> for the analytical framework.</p>
<h2>Research vs. Product Commit Patterns</h2>
<p>AI startups go through a distinctive phase transition that is visible in GitHub data:</p>
<p>Research phase: Sporadic large commits, Jupyter notebooks, experiment tracking, model checkpoints. Commit messages reference papers and experiments rather than features and fixes. Velocity is unpredictable.</p>
<p>Product phase: Frequent small commits, API endpoints, deployment configuration, monitoring setup. Commit messages reference users, features, and bugs. Velocity is sustained and accelerating.</p>
<p>The transition from research to product is the signal. When an AI startup's commit pattern shifts from sporadic-and-large to frequent-and-small, the team is moving from &quot;does this work?&quot; to &quot;let's ship this.&quot; That transition often precedes a fundraise.</p>
<h2>What to Watch in 2026</h2>
<p>The current AI sector shows interesting signal diversity. Browse the <a href="/startups-to-watch/ai-ml-q2-2026">AI &amp; Machine Learning sector rankings</a> to see who is accelerating.</p>
<p>For a broader perspective on using alternative data for deal sourcing, see <a href="/blog/alternative-data-venture-capital">why GitHub is the most underused signal in venture capital</a>.</p>]]></content:encoded>
      <pubDate>Mon, 06 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[A Weekly Deal Sourcing Workflow Using Engineering Signals]]></title>
      <link>https://signals.gitdealflow.com/blog/deal-sourcing-workflow-weekly</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/deal-sourcing-workflow-weekly</guid>
      <description><![CDATA[A 30-minute weekly workflow for investors who want to use GitHub engineering signals for deal sourcing. Step-by-step process: check rankings, screen startups, verify signals, and build a pipeline.]]></description>
      <content:encoded><![CDATA[<p>Most investors know they should use data for deal sourcing. Few have a repeatable process for doing it. Here is a 30-minute weekly workflow using engineering signals.</p>
<h2>Why a Weekly Cadence?</h2>
<p>VC Deal Flow Signal refreshes data every Monday. Engineering acceleration is a weekly signal - commit velocity change is calculated over 14-day windows. Checking more frequently than weekly adds no new information. Checking less frequently means you miss the timing advantage.</p>
<p>Monday morning is ideal: the data is fresh, and you can add qualified leads to your pipeline before the week's meetings.</p>
<h2>The 30-Minute Workflow</h2>
<p>This process is deliberately simple. The goal is not exhaustive analysis - it is fast identification of startups worth a deeper look.</p>
<h2>Step 1: Check the Trending Page (5 minutes)</h2>
<p>Open the <a href="/trending">trending page</a> and scan the top 10 startups by commit velocity change. These are the companies showing the strongest engineering acceleration across all sectors this week.</p>
<p>Look for unfamiliar names. If a company you have never heard of appears in the top 5, that is the signal working as intended - you are seeing it before mainstream channels surface it.</p>
<h2>Step 2: Filter by Your Focus Sectors (5 minutes)</h2>
<p>Navigate to the 2-3 <a href="/">sector pages</a> that match your investment thesis. Within each sector, look for startups in the top 3 that are new to you.</p>
<h2>Step 3: Screen the Top Candidates (10 minutes)</h2>
<p>For each unfamiliar startup, spend 3-4 minutes on their GitHub. The <a href="/blog/startup-engineering-metrics-investors-should-track">5-check screening framework</a> covers what to look for.</p>
<h2>Step 4: Cross-Reference (5 minutes)</h2>
<p>Search for the company on Hacker News, Twitter, and LinkedIn. Check their careers page. The goal is to confirm the engineering signal with qualitative context.</p>
<h2>Step 5: Add to Pipeline (5 minutes)</h2>
<p>For startups that pass all checks, add them to your deal tracking system. Include the engineering data: velocity change, signal type, contributor count, and the date you first noticed them. This creates a record of your timing advantage.</p>
<p>Subscribe to the <a href="https://gitdealflow.com/#signup">Signal Digest</a> to get the highlights delivered to your inbox every Monday.</p>]]></content:encoded>
      <pubDate>Sun, 05 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[5 GitHub Patterns That Predict Startup Fundraises]]></title>
      <link>https://signals.gitdealflow.com/blog/5-github-patterns-that-predict-fundraises</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/5-github-patterns-that-predict-fundraises</guid>
      <description><![CDATA[Five specific GitHub engineering patterns that have historically preceded startup fundraise announcements by 6-12 weeks. What to look for and why these patterns work as leading indicators.]]></description>
      <content:encoded><![CDATA[<p>After tracking GitHub engineering activity across thousands of startups, we have identified five patterns that consistently appear before fundraise announcements. These patterns are not guarantees, but they appear with enough regularity to be useful as leading indicators for investors. If you are new to this approach, start with our primer on <a href="/blog/how-to-read-github-signals-for-startup-investing">how to read GitHub signals for startup investing</a>.</p>
<h2>Pattern 1: What Does a Sudden Contributor Jump Signal?</h2>
<p>The most reliable fundraise predictor is a sudden, sustained increase in unique contributors. Not a gradual climb - a step function. The team goes from 5 contributors to 12 in a two-week window.</p>
<p>Why it works: most startups hire in bursts immediately after closing a round. The new hires start committing code within days of joining. If you see the contributor count jump, the round likely closed 2-4 weeks ago and the announcement is 4-8 weeks away.</p>
<p>What to look for: contributor count increases 50% or more in a 14-day window, sustained for at least 4 weeks after.</p>
<h2>Pattern 2: What Does a Burst of New Repositories Mean?</h2>
<p>A startup that suddenly creates 3-5 new public repositories in a single month is building platform infrastructure. This pattern typically appears at the Seed-to-Series-A transition: the core product works, and now the team is building the supporting ecosystem.</p>
<p>Why it works: infrastructure buildout requires capital. Companies do not invest in platform engineering unless they have runway. The timing suggests a recent or imminent fundraise.</p>
<p>What to look for: 3 or more new repositories created in 30 days, with the new repos being infrastructure-related (SDKs, APIs, internal tools, deployment configs) rather than experimental or documentation repos.</p>
<h2>Pattern 3: What Does Weekend Commit Activity Indicate?</h2>
<p>When a startup's commit pattern shifts from weekday-only to seven-days-a-week, something has changed. This is especially meaningful when the weekend activity comes from multiple contributors, not just a solo founder.</p>
<p>Why it works: teams work weekends when they are racing toward a deadline. Common triggers include a product launch, a fundraise-related demo, or a competitive response. All of these are signals that something significant is happening.</p>
<p>What to look for: sustained weekend commit activity across 2 or more contributors for 3 or more consecutive weekends.</p>
<h2>Pattern 4: Why Is a Documentation Sprint a Fundraise Signal?</h2>
<p>A sudden burst of documentation commits - README updates, API docs, architecture diagrams, contributing guides - often precedes a fundraise or launch. This is the team preparing for scrutiny.</p>
<p>Why it works: documentation is the last thing engineering teams do voluntarily. When they document proactively, they are either preparing for due diligence (fundraise), opening up to community contributions (launch), or onboarding new hires (post-fundraise). All three are interesting to investors.</p>
<p>What to look for: a week or more of documentation-heavy commits after a period of feature development. The sequence matters: code first, docs second suggests intentional preparation.</p>
<h2>Pattern 5: What Is a Velocity Regime Change?</h2>
<p>The strongest signal is not high velocity - it is a change in velocity regime. A startup that averages 30 commits per 14-day window for six months, then suddenly jumps to 90 commits for three consecutive windows, has undergone a fundamental shift.</p>
<p>Why it works: velocity regime changes reflect organizational changes. Common causes include new funding (more engineers), product-market fit (faster iteration), or a strategic pivot (rebuilding). Regime changes that sustain for 6 or more weeks are particularly meaningful.</p>
<p>What to look for: commit velocity that exceeds the 6-month average by 100% or more, sustained for 3 or more consecutive 14-day windows.</p>
<h2>How Should Investors Combine These Patterns?</h2>
<p>The patterns above are most powerful in combination. A startup showing Pattern 1 (contributor jump) and Pattern 5 (velocity regime change) simultaneously is almost certainly in the middle of a fundraise or has just closed one.</p>
<p>VC Deal Flow Signal tracks all five patterns across 15 startup sectors and classifies them into four signal types: engineering hiring burst, infrastructure buildout, deploy frequency spike, and framework migration. For the full metrics checklist, see <a href="/blog/startup-engineering-metrics-investors-should-track">7 engineering metrics every investor should track</a>. Browse the sector rankings to see which startups are showing these patterns right now.</p>]]></content:encoded>
      <pubDate>Sat, 04 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Cybersecurity Startup Signals: Reading GitHub Data for Security Deals]]></title>
      <link>https://signals.gitdealflow.com/blog/cybersecurity-startup-signals</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/cybersecurity-startup-signals</guid>
      <description><![CDATA[Cybersecurity startups have unique GitHub patterns: rapid response to CVEs, compliance-driven sprints, and infrastructure hardening. Learn what cybersecurity engineering signals mean for investors.]]></description>
      <content:encoded><![CDATA[<p>Cybersecurity startups present unique challenges for GitHub-based signal analysis. The sector's engineering patterns are driven by threat response cycles and compliance requirements in ways that other sectors are not.</p>
<h2>CVE-Driven Development</h2>
<p>The most distinctive cybersecurity pattern: deploy frequency spikes that correlate with CVE disclosures. When a major vulnerability is published, security companies rush to patch, update, and ship. This creates commit velocity spikes that are reactive, not strategic.</p>
<p>For investors, the question is whether a velocity spike reflects incident response or product momentum. Check the timing: does the spike coincide with a major CVE disclosure? If so, the acceleration is defensive, not offensive.</p>
<h2>Compliance Infrastructure Signals</h2>
<p>Like fintech, cybersecurity startups build significant compliance infrastructure: SOC 2 audit trails, ISO 27001 documentation, penetration testing frameworks, and security certification tooling.</p>
<p>New repositories related to compliance indicate a company preparing for enterprise sales - most enterprise buyers require SOC 2 compliance at minimum. This is a positive investment signal because enterprise-readiness requires capital and precedes revenue growth.</p>
<h2>The Strongest Cybersecurity Signal</h2>
<p>The most compelling cybersecurity investment signal is sustained engineering acceleration that is not correlated with external events. When a security startup is shipping fast without a CVE trigger or compliance deadline, the team is building something new. That organic acceleration is the same signal that works across all sectors - and it precedes fundraising by the same 6-12 week window.</p>
<p>Browse the <a href="/startups-to-watch/cybersecurity-q2-2026">Cybersecurity sector rankings</a> to see which security startups are showing engineering acceleration right now.</p>]]></content:encoded>
      <pubDate>Sat, 04 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[Climate Tech Engineering Signals: What GitHub Data Reveals About Green Startups]]></title>
      <link>https://signals.gitdealflow.com/blog/climate-tech-engineering-signals</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/climate-tech-engineering-signals</guid>
      <description><![CDATA[Climate tech startups combine hardware and software development, creating unique GitHub patterns. Learn how to interpret engineering signals for energy, carbon, and sustainability startups.]]></description>
      <content:encoded><![CDATA[<p>Climate tech is one of the fastest-growing sectors in venture capital. But it also one of the hardest to analyze with software engineering metrics, because many climate tech companies build physical products.</p>
<h2>The Software-Hardware Spectrum</h2>
<p>Climate tech startups fall along a spectrum from pure software to pure hardware. At the software end: carbon accounting platforms, renewable energy trading tools, grid optimization software, and ESG reporting systems. These companies look like enterprise SaaS on GitHub - high commit velocity, standard signal patterns.</p>
<p>At the hardware end: battery manufacturers, solar panel companies, and industrial decarbonization. These companies may have minimal public GitHub activity because most of their engineering work is in hardware design, not software.</p>
<p>The most interesting companies for GitHub signal analysis sit in the middle: hardware-adjacent software companies that build the intelligence layer for physical systems. Battery management systems, sensor networks, predictive maintenance for wind farms, and energy grid optimization.</p>
<h2>What Climate Tech Signals Look Like</h2>
<p>Software-heavy climate tech: Looks like enterprise SaaS. Commit velocity, contributor growth, and signal types follow standard patterns. Use the same analytical framework as any other sector.</p>
<p>Hardware-adjacent climate tech: Lower absolute commit velocity, but meaningful infrastructure buildout signals. New repositories for data pipelines, IoT integrations, and edge computing indicate a transition from R&amp;D to deployment.</p>
<h2>The R&amp;D-to-Deployment Transition</h2>
<p>The strongest climate tech investment signal is the R&amp;D-to-deployment transition. When a company's GitHub activity shifts from experimental (research notebooks, prototype code) to operational (deployment scripts, monitoring, CI/CD), the technology is moving from lab to field.</p>
<p>This transition requires capital - deploying physical systems costs money. Engineering acceleration during this phase is a strong fundraise predictor.</p>
<p>Browse the <a href="/startups-to-watch/climate-tech-q2-2026">Climate Tech sector rankings</a> to see which green startups are showing acceleration right now.</p>]]></content:encoded>
      <pubDate>Fri, 03 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[5 Mistakes Investors Make When Reading GitHub Signals]]></title>
      <link>https://signals.gitdealflow.com/blog/investor-mistakes-github-signals</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/investor-mistakes-github-signals</guid>
      <description><![CDATA[Common pitfalls when using GitHub engineering data for deal sourcing: confusing stars with traction, ignoring private repos, overweighting absolute velocity, missing the context behind spikes, and treating signals as investment decisions.]]></description>
      <content:encoded><![CDATA[<p>GitHub engineering signals are a powerful deal sourcing tool. They are also easy to misread. Here are the five most common mistakes investors make - and how to avoid them.</p>
<h2>Mistake 1: Confusing Stars with Traction</h2>
<p>GitHub stars measure social interest. They are a vanity metric, not an engineering signal. A repository with 10,000 stars may have zero commercial traction. A repository with 50 stars may power a company with $5M ARR.</p>
<p>Stars tell you what developers find interesting. Commit velocity tells you what companies are actually building. Focus on the latter.</p>
<h2>Mistake 2: Ignoring the Private Repo Blind Spot</h2>
<p>Public GitHub activity is a biased sample. Many startups - especially in enterprise SaaS, fintech, and healthcare - keep all their code in private repositories. A startup with no public GitHub activity is not necessarily inactive; it may be very active in ways you cannot see.</p>
<p>This means GitHub signals work best for sectors with a culture of open source or public development: developer tools, infrastructure, AI/ML, and Web3. For sectors with strong privacy norms, use GitHub signals as one input among many.</p>
<h2>Mistake 3: Overweighting Absolute Velocity</h2>
<p>A startup with 500 commits per week is not necessarily more interesting than one with 50. What matters is the rate of change - is the 50-commit startup accelerating?</p>
<p>Absolute velocity correlates with team size, not with momentum. <a href="/glossary#commit-velocity-change">Commit velocity change</a> normalizes for team size by measuring acceleration relative to the company's own baseline.</p>
<h2>Mistake 4: Missing Spike Context</h2>
<p>Not all velocity spikes are positive signals. Common false positives:
- Documentation sprints (high commit count, low engineering substance)
- CI/CD bot activity (automated commits inflating counts)
- One-time migrations (framework upgrades, monorepo restructuring)
- Hackathon artifacts (intense activity that does not sustain)</p>
<p>The fix: spend 5 minutes on the GitHub organization before acting on a spike. Check recent commit messages and which repositories are active. This is step 3 in our <a href="/blog/deal-sourcing-workflow-weekly">weekly sourcing workflow</a>.</p>
<h2>Mistake 5: Treating Signals as Decisions</h2>
<p>Engineering acceleration is a sourcing signal, not an investment thesis. It tells you which companies are worth investigating - not which companies are worth investing in.</p>
<p>The signal gets you to the table early. The decision still requires founder conversations, product evaluation, market analysis, and competitive landscape assessment. GitHub data gives you timing advantage; due diligence gives you conviction.</p>
<p>For the full screening framework, see the <a href="/blog/startup-engineering-metrics-investors-should-track">7 engineering metrics every investor should track</a>.</p>]]></content:encoded>
      <pubDate>Thu, 02 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[How VCs Use GitHub for Technical Due Diligence]]></title>
      <link>https://signals.gitdealflow.com/blog/github-due-diligence-for-vcs</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/github-due-diligence-for-vcs</guid>
      <description><![CDATA[A practical framework for using public GitHub data in venture capital due diligence. What to look for, what to ignore, and how engineering signals complement traditional diligence methods.]]></description>
      <content:encoded><![CDATA[<p>Technical due diligence is one of the most time-consuming parts of the venture investment process. VCs typically hire external consultants, schedule deep-dive sessions with engineering teams, and review architecture documents. This process takes weeks and often happens late in the deal cycle.</p>
<p>Public GitHub data cannot replace a proper technical deep dive. But it can do something equally valuable: help you decide which companies deserve that deep dive in the first place.</p>
<h2>What Can Public GitHub Data Tell Investors?</h2>
<p>GitHub profiles reveal several dimensions of engineering health that are useful for investors:</p>
<p>Engineering velocity and consistency: Is the team shipping regularly, or are there long gaps followed by frantic bursts? Consistent commit patterns suggest disciplined engineering practices. Erratic patterns may indicate management instability, pivots, or part-time teams.</p>
<p>Team composition signals: Contributor counts, contribution patterns, and the ratio of organizational contributors to external ones reveal team structure. A startup with 3 contributors making 90% of commits has a different risk profile than one with 15 active contributors.</p>
<p>Technology choices: The programming languages, frameworks, and tools visible in public repositories tell you about technical maturity. A seed-stage startup using enterprise-grade infrastructure tooling may be over-engineering. A growth-stage company still on prototype-quality tools may have technical debt.</p>
<p>Open source strategy: Some startups use open source as a go-to-market channel (developer tools, infrastructure). Their GitHub activity IS the product signal. Others keep everything private and only have minor utility repos public. The absence of public activity is not a negative signal for the latter.</p>
<h2>What Are the Limitations of GitHub Data for Due Diligence?</h2>
<p>Code quality: Commit volume says nothing about code quality, test coverage, or architectural soundness. A team making 200 commits a week could be writing excellent code or terrible code.</p>
<p>Private repository activity: Most startups keep their core product code private. Public repos may represent only a fraction of actual engineering work. Never assume low public activity means low engineering output.</p>
<p>Individual contributor value: Not all contributors are equal. One senior engineer making 10 thoughtful commits may contribute more value than five junior developers making 50 commits each.</p>
<p>Business context: Engineering acceleration without business context is just a number. The same commit pattern could indicate product-market fit, a desperate pivot, or a hackathon project.</p>
<h2>How Should Investors Use GitHub in Their Due Diligence Process?</h2>
<p>Here is how to use GitHub data at each stage of the investment process:</p>
<p>Sourcing stage: Use commit velocity change to identify startups worth researching. This is what VC Deal Flow Signal automates - surfacing the companies showing unusual engineering acceleration. See the <a href="/blog/startup-engineering-metrics-investors-should-track">7 engineering metrics every investor should track</a> for a complete checklist.</p>
<p>Initial screening: Look at the GitHub organization profile. How many public repos? When was the last push? Is there a pattern of consistent activity, or sporadic bursts? This takes 2 minutes and can save you from scheduling calls with inactive teams.</p>
<p>Pre-meeting research: Before a founder meeting, check their GitHub. What languages and frameworks do they use? How many active contributors? This gives you informed questions to ask during the call.</p>
<p>Post-meeting verification: After hearing the founder's story about their engineering team and roadmap, cross-reference with GitHub. Does the team size they claimed match contributor counts? Does their claimed velocity match commit patterns?</p>
<p>Portfolio monitoring: After investing, use GitHub signals as an early warning system. A portfolio company whose commit velocity drops 50% over two months may be experiencing team attrition, strategic confusion, or runway pressure. This signal appears before the quarterly board update.</p>
<h2>What Are the Ethical Considerations of Using GitHub Data?</h2>
<p>Using public data for investment decisions is legal and common. However, there are ethical considerations:</p>
<p>Do not contact individual contributors or attempt to recruit from portfolio companies. GitHub profiles are public, but using them to poach talent is poor form in the investor community.</p>
<p>Do not make investment decisions based solely on GitHub data. It is one signal among many. The strongest investment thesis combines engineering signals with market analysis, founder evaluation, and customer reference checks.</p>
<p>Always remember that engineering acceleration, as VC Deal Flow Signal defines it, measured from public GitHub commit and contributor activity, is a leading indicator, not a guarantee. Some of the fastest-accelerating startups will fail. The data gives you timing advantage, not outcome certainty. For the specific patterns to watch, read about the <a href="/blog/5-github-patterns-that-predict-fundraises">5 GitHub patterns that predict fundraises</a>.</p>]]></content:encoded>
      <pubDate>Wed, 01 Apr 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[How to Read GitHub Signals for Startup Investing]]></title>
      <link>https://signals.gitdealflow.com/blog/how-to-read-github-signals-for-startup-investing</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/how-to-read-github-signals-for-startup-investing</guid>
      <description><![CDATA[A practical guide for investors on interpreting GitHub engineering activity as a leading indicator of startup traction. Covers commit velocity, contributor growth, and what patterns actually predict fundraises.]]></description>
      <content:encoded><![CDATA[<p>GitHub is the largest free dataset of real-time engineering activity in the world. Every public commit, every new repository, every contributor who joins a project - it is all timestamped and queryable. Yet almost no investor uses it for deal sourcing.</p>
<p>The reason is simple: raw GitHub data is noisy. Thousands of commits a day across millions of repositories. Without a framework for what matters, it is just noise.</p>
<p>This post explains the framework we use at VC Deal Flow Signal to turn GitHub activity into actionable deal flow intelligence.</p>
<h2>What Is Engineering Acceleration?</h2>
<p>We do not measure absolute engineering output. A company with 500 commits a week is not necessarily more interesting than one with 50. What matters is the rate of change - acceleration.</p>
<p>When a startup's commit velocity doubles in two weeks, something has changed. Maybe they just closed a seed round and are shipping furiously. Maybe they hired three engineers and are building out infrastructure. Maybe they found product-market fit and are iterating fast on customer feedback.</p>
<p>Whatever the cause, the effect is visible in the commit graph weeks before it appears in a press release or a pitch deck landing in your inbox. We have identified <a href="/blog/5-github-patterns-that-predict-fundraises">five specific GitHub patterns that predict fundraises</a> with the most consistency.</p>
<h2>What Are the Four Types of Engineering Signals?</h2>
<p>We classify engineering acceleration into four patterns:</p>
<p>Engineering hiring burst: Contributor count jumps 50% or more in a short window. This usually means the company just closed a round and is scaling the team. If you are seeing this signal, you are likely too late for the current round - but perfectly timed for the next one.</p>
<p>Infrastructure buildout: Three or more new public repositories created in 30 days. The company is expanding its technical surface area - new microservices, new SDKs, new internal tools. This is classic Series A behavior: the product works, now they are building the platform.</p>
<p>Deploy frequency spike: Commit velocity increases 150% or more versus baseline. The team is shipping at an unusually high rate. This can indicate a product launch, a pivot, or a response to sudden customer demand. All are interesting to investors.</p>
<p>Framework migration: General acceleration that does not fit the above categories. Often indicates a technology stack transition - moving from a prototype stack to a production stack. This is the subtlest signal but can indicate the shift from exploration to exploitation.</p>
<h2>What GitHub Activity Is Not a Useful Signal?</h2>
<p>Not all GitHub activity is meaningful for investors:</p>
<p>- Open source maintenance: Popular open source projects have high commit volumes but that tells you nothing about the company's product trajectory.
- Documentation pushes: A burst of markdown commits usually means a docs sprint, not product acceleration.
- CI/CD noise: Some teams commit generated files or configuration changes that inflate commit counts without reflecting product work.</p>
<p>We mitigate these by measuring change from baseline rather than absolute counts. A docs sprint looks different from a product sprint when you compare the commit graph to the company's own history.</p>
<h2>When Do Engineering Signals Appear Before Fundraises?</h2>
<p>In our data, engineering acceleration signals precede fundraise announcements by three to six weeks on average. The pattern looks like this:</p>
<ol><li>Weeks 1-2: Commit velocity starts climbing. Contributor count may tick up.</li><li>Weeks 3-4: Acceleration becomes obvious. New repositories appear. Signal type becomes classifiable.</li><li>Weeks 5-8: The company is heads-down building. If they are raising, the round is in progress but not yet announced.</li><li>Weeks 8-12: Fundraise announcement, TechCrunch article, your inbox lights up with the same deck everyone else got.</li></ol>
<p>If you are reaching out in weeks 2-4, you are ahead of the crowd. That is the window this data gives you.</p>
<h2>How Should Investors Use This in Practice?</h2>
<p>The most effective approach is sector-focused. Pick two or three sectors you know well and watch the weekly rankings:</p>
<ol><li>When a startup you do not recognize appears in the top 3, research them.</li><li>Look at their GitHub: is the activity product-related or infrastructure-related?</li><li>Cross-reference with Crunchbase: are they pre-raise? Post-raise and scaling?</li><li>If the signal is strong and the timing is right, reach out to the founder.</li></ol>
<p>The worst thing you can do with this data is use it as a replacement for judgment. Engineering acceleration is a leading indicator, not a guarantee. But combined with sector expertise and founder evaluation, it gives you a structural timing advantage that most investors do not have. For a deeper look at technical evaluation, see our guide on <a href="/blog/github-due-diligence-for-vcs">how VCs use GitHub for due diligence</a>.</p>
<h2>Where Can I Start Watching?</h2>
<p>We track engineering acceleration across 15 startup sectors, updated weekly. Each sector page ranks the top startups by commit velocity change and classifies their signal type.</p>
<p>Browse the sector rankings to see which startups are accelerating right now.</p>]]></content:encoded>
      <pubDate>Sat, 28 Mar 2026 00:00:00 GMT</pubDate>
    </item>
    <item>
      <title><![CDATA[What Is Deal Flow Signal? 4 Signal Types, 6-12 Week Lead Time (2026)]]></title>
      <link>https://signals.gitdealflow.com/blog/what-is-deal-flow-signal</link>
      <guid isPermaLink="true">https://signals.gitdealflow.com/blog/what-is-deal-flow-signal</guid>
      <description><![CDATA[Deal flow signal refers to data-driven indicators that help investors identify promising startups before traditional channels surface them. Learn how engineering momentum serves as a leading indicator of traction.]]></description>
      <content:encoded><![CDATA[<p>A deal flow signal is any data-driven indicator that helps an investor identify a promising startup before traditional channels surface it. Traditional deal flow relies on warm introductions, pitch decks, and press coverage; signals supplement it with quantitative, real-time data that appears weeks or months earlier.</p>
<h2>What Are the Main Types of Deal Flow Signal?</h2>
<p>There are four main types of deal flow signal, each with a different lead time before a fundraise announcement:</p>
<ol><li>Engineering signals (6-12 weeks lead time): changes in a startup's public GitHub activity, commit velocity, contributor growth, and repository creation. Engineering acceleration precedes product milestones, which precede fundraise decisions.</li><li>Hiring signals (4-8 weeks): job postings, especially for senior engineering and go-to-market roles, indicate growth plans and budget.</li><li>Web traffic signals (4-6 weeks): rapid growth in a startup's web traffic can indicate product-market fit before revenue shows up in databases.</li><li>Social signals (1-2 weeks): mentions on Twitter, Hacker News, Product Hunt, and industry forums. By the time a startup trends on social media, most investors are already aware.</li></ol>
<h2>Why Is Traditional Deal Flow Not Enough?</h2>
<p>Most VCs source deals through their network. The problem is that networks are shared. By the time a startup is making the rounds at demo day or landing in your inbox via a warm intro, it is also landing in every other investor's inbox.</p>
<p>The result is that competitive deals - the ones most likely to generate outsized returns - are identified late and negotiated under pressure. The investor who arrives first has a structural advantage: they set the terms, they build the relationship before the founder is overwhelmed with options. This is why <a href="/blog/alternative-data-venture-capital">alternative data is becoming essential for venture capital</a>.</p>
<h2>Why Are GitHub Signals the Best Leading Indicator?</h2>
<p>GitHub engineering activity has unique properties that make it the most reliable early deal flow signal:</p>
<ol><li>It is hard to fake. Commits represent actual work. You cannot game commit velocity the way you can game social media metrics.</li><li>It is continuous. Unlike hiring signals (which appear when a job is posted) or press (which appears when a company wants attention), engineering activity happens daily.</li><li>It is free and public. Unlike web traffic data (which requires third-party tools) or hiring data (which requires scraping job boards), GitHub data is available via a public API.</li><li>It reveals intent. The type of engineering work - infrastructure buildout vs. feature shipping vs. team scaling - tells you what phase the company is in.</li></ol>
<h2>How Does VC Deal Flow Signal Work?</h2>
<p>We monitor GitHub engineering activity across 15 startup sectors. For each sector, we:</p>
<ol><li>Identify active startup organizations using topic-based search.</li><li>Pull commit activity, contributor data, and repository creation data.</li><li>Calculate 14-day commit velocity and its rate of change.</li><li>Classify the signal type (hiring burst, infrastructure buildout, deploy spike, framework migration).</li><li>Rank startups by engineering acceleration (GitHub commit velocity + contributor growth, not accelerator-program participation).</li></ol>
<p>The result is a weekly-updated ranking of startups showing the strongest engineering momentum in each sector. Investors can use this to identify breakout companies weeks before they appear through traditional channels.</p>
<h2>How Do I Get Started with Deal Flow Signal?</h2>
<p>The simplest way to start using deal flow signal is to get our free Signal Report - five breakout startups with real GitHub acceleration data, delivered weekly. For deeper access, our Dashboard gives you the full ranked list across all 15 sectors with filtering by stage, geography, and signal type. To learn the practical framework, read our guide on <a href="/blog/how-to-read-github-signals-for-startup-investing">how to read GitHub signals for startup investing</a>.</p>]]></content:encoded>
      <pubDate>Wed, 25 Mar 2026 00:00:00 GMT</pubDate>
    </item>
  </channel>
</rss>