A dual Mac Mini cluster running 183+ AI skills, multi-model orchestration, autonomous software pipelines, financial analysis engines, marketing engine, LLM council, and a self-healing infrastructure — replacing over $250,000/year in SaaS tools and human coordination.
The numbers behind a two-machine system that thinks, codes, trades, researches, markets, and manages itself.
Two specialized machines, directly linked via Thunderbolt 4, forming one coherent autonomous operation with hot standby and auto-failover.
Each pillar is a self-contained capability domain. Together, they form a system that operates, decides, creates, markets, and researches — independently.
Six cloud models routed by task type and pipeline, enforcing model diversity to prevent single-provider lock-in.
Lightweight local inference on Forge — always-on embeddings, on-demand testing.
Heavy LLM inference on Anvil's 48GB M4 Pro — daily tasks, vision, code generation.
Two pipelines — emergE is the default for ALL projects (Fulcrum AI, Nexdex, VerySmart, new ventures). Vibestreet keeps its own tuned pipeline until it ships:
The default for all new projects. 90% cheaper than v1. Model diversity: DeepSeek → GLM → DeepSeek → Anthropic → Anthropic.
End-to-end workflow from idea to deployed feature — now with Architecture Review Gate for emergE:
Composable engineering sub-skills loaded on demand:
Full SEC filing extraction and 3-statement modeling.
Quantitative trading research and model development.
Semantic knowledge engine — PostgreSQL 17 + pgvector on Anvil, WAL-streamed to Forge standby.
6 Obsidian vaults organized by venture and domain:
Different persistence guarantees for different needs:
Bridges session memory and permanent storage — no context lost across restarts.
Multi-provider routing to OpenAI GPT-Image, Fal.ai Flux, Google, and more.
Text-to-video, image-to-video, and video-to-video with multi-modal references.
Multi-provider audio creation including Google Lyria.
Configured vision model for inspection and understanding — qwen3-vl:8b always-on on Anvil.
Blocks dangerous operations before they execute.
Full audit trail after every tool execution.
18-pattern scanner prevents secrets from entering version control.
Continuous monitoring with automatic failover — not just restart.
Venture-scoped channels for focused operational context.
Time-sensitive alerts and briefings to leadership.
Browser-based administration and monitoring.
Wake-word activated voice assistant.
Scoped permissions for team collaboration.
Cron-driven automation for self-maintaining operations.
Anti-detection web automation with full session control.
Dedicated agent for project management and issue processing.
3-generation rolling backup with multiple storage tiers.
Two Apple Silicon Mac Minis — energy-efficient, silent, neural engine on-chip.
Core operations have zero external dependencies.
22 LaunchDaemons on Forge — all auto-restart capable.
5 LaunchDaemons on Anvil — compute-focused, self-contained.
25-layer architecture specification for edge AI compute nodes.
Capture skill. Converts raw input into a structured brief. Preserves energy, quotables, and intent.
9-stage orchestration graph that chains all other marketing skills into a coherent campaign.
Explicit model assignment per content type. No one model writes everything.
7-check quality gate. A different model evaluates than the one that drafted. Score below 56 = blocked.
Closes the loop. Every campaign gets measured and fed back into the next one.
End-to-end campaign execution:
Multi-source data collection, synthesis, and structured output for any topic.
Scans incoming data streams for actionable patterns before they become obvious.
Autonomous web research agent. Given a topic, it browses, extracts, structures, and stores knowledge.
Deep competitor analysis producing structured battle cards for sales and positioning.
Multi-model ensemble validation for trading signals. No signal acts without cross-model agreement.
High-stakes decisions deserve more than one opinion. LLM Council fans queries to 3-5 models, collects independent positions, then synthesizes a final answer.
Direct access to original academic papers and textbooks — no more relying on secondhand blog summaries.
What it actually replaces — in hard numbers.
Total monthly operating cost for the entire dual-machine cluster — less than a single SaaS subscription. DeepSeek V4-Flash reduced cloud AI spend by ~15%.
How a single request flows through the two-machine cluster — end to end, autonomously.
Chairman sends a message via Discord, WhatsApp, or voice. Forge Gateway receives it and routes to the primary AI model with full context loaded — memory, entity state, active threads. Forge handles all inbound channel routing.
Agent scans 183 skills. If one matches, it loads automatically. If the task is complex, a goal is created and a multi-step plan is generated. Sub-agents may be spawned for parallel work. Light tasks stay on Forge; heavy compute is delegated to Anvil.
Light tasks (routing, API calls, web research, file ops) execute on Forge. Heavy tasks (LLM inference, code builds, simulations) are delegated to Anvil via Thunderbolt 4. The security layer checks every tool call pre-execution and logs every result post-execution on both machines.
Every decision, output, and insight is captured. Memory files written on Forge are rsync'd to Anvil every 5 minutes. Knowledge graph updates go to Anvil's PostgreSQL primary and WAL-stream to Forge standby. State persistence layer tracks all entity changes across both machines.
Response delivered via originating channel with platform-appropriate formatting. Background monitoring continues — health checks every 2 minutes across both machines, backups on schedule, auto-failover if Anvil goes down, auto-recovery if any service crashes. The cluster never sleeps.