Pipeline, warehouse, modern-data-stack, ML-engineering practitioners · 12 rooms · ranked by yield
Data and ML engineering collectives. Ranked by deal-flow yield.
dbt, MLOps, r/dataengineering, Modern Data Stack, where production-data orgs first leak hiring + roadmap.
Data-eng and ML-eng practitioner rooms. The audience overlaps with /infra-and-platform-eng-circles but indexes harder on data-platform and modern-data-stack adoption. Crossover with /ai-builder-rooms is intentional, same people, different lens.
Yield story
Primary surfaces are dbt Community Slack, MLOps Community Slack, r/dataengineering, r/MachineLearning, DataTalks.Club Slack, and Modern Data Stack forum, six rooms where data-platform repos most often first telegraph commercial plans.
Cadence
Weekly read on the six primary surfaces.
Caveat
Data-eng rooms reward operational precision. No marketing-shaped posts.
Yield breakdown · 12 rooms
6
Primary
Direct deal-flow surface. Multiple repos we surface first showed up here.
5
Secondary
Adjacent. Signals overlap, but rarely the first-mention surface.
1
Ambient
Culture / reading list. Useful for framing, not for sourcing.
Engagement status
0
Engage
8
Watch
0
Hold
4
Read
0
Blocked
The ranked roster
12 rooms, sorted by yield tier, numbered, linked.
Primary first, then secondary, then ambient. Hand-authored order within each tier, top of tier = highest attention density.
- 1.
dbt-practitioner Slack. Modern-data-stack adoption signal.
- 2.
Production-ML practitioners. Crossover with AI-builder rooms.
- 3.
Data-platform stack threads. High GitHub-org overlap.
- 4.
Where ML papers and infrastructure orgs first surface.
- 5.
Data-eng practitioner Slack. Cohort-of-engineers signal.
- 6.
MDS vendor + buyer crossover surface.
- 7.
Analytics-practitioner Slack. Niche but high-signal-density.
- 8.
Hex-customer notebook community.
- 9.
Data-science practitioner audience.
- 10.
Experiment-tracking customers. Research-to-product flips.
- 11.
HF discussion forum. Long-tail of model-product transitions.
- 12.
Snowflake-power-user community. Adjacent buyer pool.
Frequently asked
Four questions about this list.
- What is the "deal-flow yield" of a community?
- Yield is a three-tier rating, primary, secondary, ambient, describing how often a community has been the first public surface where a repo or founder we now surface inside the product appeared. Primary means multiple deals trace back to the room; ambient means the room is read for framing, not sourcing.
- Why is data and ml engineering collectives indexed by type instead of by platform?
- The platform map at /voices already covers per-platform rosters (Reddit-100, HN-100, X-100). Indexing by type cuts the same audience differently, a Cursor user in a Cursor Discord is a different deal-flow signal than the same user in a CNCF channel, even though both are on Discord.
- Do we engage in every community on this page?
- No. Each community is tagged with an engagement status (engage / watch / hold / read / blocked). The default for most rooms is "watch" or "read", we listen but do not post. The "engage" tag is reserved for a small number of rooms where comment-only engagement on technical Q&A is allowed by both the room rules and our anonymity rule.
- Are the individuals inside these communities listed anywhere?
- No. The anonymity rule that governs the entire site applies here: we name communities, not individuals. The specific founders, maintainers, and contributors we surface live inside the paid product, that is the buyer's edge.
Where the deal-flow signal actually lives
Rooms tell us where to look. Telemetry tells us what to fund.
Reading data and ml engineering collectivesdoesn’t produce a deal. It sharpens the question we ask the GitHub-side telemetry. The actual deal-flow signal lives at /predicted, /firstlook, and the rest of the signal stack below.
Data and ML engineering collectives roster regenerated quarterly. Status flags update weekly. Companion platform map at /voices. Master mixed list at /target-list.