📅

Prompt-Intelligenz — 24. September 2026

🏆 Highlight

### „Voice Design per Prompt" — Gemini 3.8 Flash TTS macht Stimmen zu einer Natural-Language-Bauanleitung Prompt (vollständig, kopierbar — Bausteine aus Googles Dokumentation und Demo-Galerie): ``` Voice: a high-energy DJ voice from Melbourne Voice: a super-tinny, monotone robot voice Voice: a dramatic, fire-breathing dragon with a distinct regional cadence (Japanese) Remixing (Voi...

Fresh HN data loaded. Today's interesting stories: Gemini 3.8 TTS, Contrastive Language Models, rogue AI agents. Let me fetch front page, story details, and other sources in parallel. Front page analyzed. Key prompt-relevant stories: Gemini 3.8 TTS, Claude enzyme discovery, LensVLM, Contrastive LMs. Now fetching article bodies + HN comment threads in parallel. Simon Willison has a new Gemini TTS playground post (found via HN comment). Fetching it plus checking enzyme/contrastive comment threads. PDF says system prompt not included — no exact prompt there. Moving to GitHub trending, news portals, and arXiv in parallel. Great GitHub finds: 3dicon (272★, Claude Code skill), ComfyUI-QwenImage-PhotoStyles, awesome-opus-5-5-prompts. Fetching READMEs + parsing news portals. Fetching key MarkTechPost articles (Gemini TTS voice prompts, Qwen-Image-2.1, Jev guide) and reading the GitHub READMEs. 3dicon and prompt-fossils are strong. opus-55-10k-websites has full reusable prompts in /prompts/. Fetching those + parsing MarkTechPost articles. Strong material gathered. Now checking previous report format, fetching Simon's TTS playground tool, and reading the drone prompt's video section. Format confirmed. One final quick scan (Opus prompts repo, Jev guide, enzyme quote, Google TTS prompt fragments), then writing the report. All research done. Extracting the full CatWalk prompt text, then writing the report immediately. Research complete (~11 min). Writing the final report now.

Prompt-Intelligenz — 24. September 2026

Die besten kopierbaren Prompts der letzten 24 Stunden, kuratiert für prompta.ch — heute zwischen Googles prompt-gesteuerter TTS-Regie (Gemini 3.8), der ersten promptsammlung für Claude Opus 5.5, einem Fossilien-Audit für CLAUDE.md-Dateien und der Frage, warum das Wort „drone" in Videoprompts verboten ist.

🔤 TOP 3 PROMPTS — Textgenerierung

1. „CatWalk" — der elegante Spielprompt für Claude Opus 5.5

Prompt (vollständig, kopierbar):

Let's develop a game.
The graphics can be either 2D or 3D; choose whichever makes the game systems and overall development easier to analyze and implement.
Personally, I'm imagining a 3D game with a side-scrolling action format.
I'd like to create a stylish atmosphere through the lighting and other elements, so I think 3D could enable more beautiful visuals—for example, streetlamps and lanterns.
The game I want to make is called CatWalk.
The name says it all.
A cat moves from side to side.
A catwalk stretches ahead, and the screen scrolls automatically. The player uses only simple controls, such as jumping, to clear obstacles and gaps while keeping pace with the scrolling speed. In a way, it might have a similar tension and system to Flappy Bird.
However, I want the graphics to feel mature, stylish, and atmospheric.
If possible, I'd love to show the cat moving, running, and jumping with graceful, fluid animation.
You can decide the stage's setting, but starting with something straightforward, such as a nighttime street, would be fine.
I'd be very happy if a darker setting could make indirect lighting and similar effects look beautiful.
I understand that some things may be possible, impossible, or difficult.
Using my requests as a starting point, please develop something you think you can realistically make.
For now, make one stage playable through a single complete loop.

Am besten mit: Claude Opus 5.5 (in Claude Code oder einem Agenten mit Datei- und Preview-Zugriff)

Warum effektiv: Der Prompt gibt ein klares Spielkonzept vor, überlässt aber Technik-Entscheidungen (2D/3D) ausdrücklich dem Modell — „choose whichever makes the game systems easier to analyze and implement". Wünsche werden als Prioritäten formuliert („mature, stylish, and atmospheric"), nicht als Befehle, und die realistische Erwartungsschleuse am Ende („some things may be possible, impossible, or difficult") verhindert Über-Versprechen. Der Scope ist auf eine spielbare Loop begrenzt — genau die richtige Größe für einen ersten Wurf.

Quelle: https://github.com/TripoGrowthLab/awesome-opus-5-5-prompts | 4 GitHub-Sterne (Repo seit gestern neu)

Community Resonanz: Die Sammlung sammelt source-verlinkte Opus-5.5-Prompts für 3D-Szenen, Spiele, Animationen und Simulationen und ist bereits in 14 Sprachen übersetzt — die Community baut sie als Referenz für die neue Modellgeneration auf.

2. „Cinematic Drone Fly-Through" — die komplette Website-Bauanweisung

Prompt (vollständig, kopierbar):

Act as a designer, creative director and website developer. Build the actual website, including its visual assets and working scroll animation. The centrepiece is one continuous first-person fly-through of the business: a single unbroken camera move that enters the building, travels through its spaces the way a real FPV pilot would fly them, exits, and reveals the whole place from the air. Motion, typography and content should feel composed together.

Use Higgsfield MCP for generated images and video. Stay in the current conversation and project. Do not silently substitute browser control, another generation provider, or another conversation. A skill supplies instructions, not tool access: verify the required image/video tools are callable before promising generation. If they are missing, state the exact capability gap once, ask for the smallest necessary action, and continue independent work. Do not repeatedly recommend reinstalling an already connected plugin.

## 1. Establish the brief with minimal friction

Read the conversation and supplied materials first. Ask only questions whose answers materially affect the result. When information is missing, a single compact intake can cover:

- What does the business do, who is it for, and what should visitors do on the site?
- Is there a business name, logo or existing brand identity to use?
- Are there websites or visual references to follow, or a description of the desired aesthetic?
- What is the physical route? Which spaces should the camera pass through (for example entrance, showroom, workshop, storage, rear yard), and where should it exit?

Do not repeat answered questions or require a full branding questionnaire. A business description is enough to begin; if the route is not given, propose a realistic one for that type of premises.

Recognize two modes:

Guided mode: Use a small number of meaningful review points: identity/style direction, the start still, and the proposed flight path. Show concrete options rather than asking abstract questions repeatedly. Once a direction is chosen, move forward.

Template or autonomous mode: If the user says "build a template," "choose for me," "assume the answers," or similar, choose suitable defaults and complete the build without waiting for aesthetic decisions. Missing copy is not a blocker. Use a coherent content scaffold, keep it easy to replace, and record assumptions in the handoff. The user can supply final business details later.

Preserve real supplied facts. Do not fabricate testimonials, customers, awards, addresses, experience, project counts or other credibility claims. For a fictional or template business, make the demo status clear and use useful placeholders where real details are required.

Am besten mit: Claude Opus 5.5 in Claude Code, mit Higgsfield MCP (Bild- und Video-Tools) verbunden

Warum effektiv: Der Prompt ist ein vollständiger Projektleiter in Textform: Er definiert zwei Betriebsmodi (Guided vs. Autonomous), verbietet Tool-Substitution und Endlos-Fragebögen und erzwingt Anti-Fabrication-Regeln für Glaubwürdigkeitsaussagen. Der Trick „A skill supplies instructions, not tool access: verify the required tools are callable before promising generation" verhindert die häufigste Agenten-Falle — das Versprechen von Generierung ohne funktionierende Tools. Die Vollversion (18 KB, inkl. Flight-Path-Design, Scroll-Animation und Handoff) liegt im Repo unter prompts/01-desktop-drone-flythrough.md.

Quelle: https://github.com/Barty-Bart/opus-55-10k-websites | 4 GitHub-Sterne

Community Resonanz: Der Autor (Bart Slodyczka) liefert zwei Production-Prompts plus Video-Walkthrough — als Vorlage für „10K-Websites", die sich per Scroll-Steuerung durchfliegen lassen; die Prompts sind direkt als .md-Dateien kopierbar.

3. „Kriegs-Atlas-Duell" — der Einzeiler, der zwei Top-Modelle gegeneinander antreten lässt

Prompt (vollständig, kopierbar):

hey, make a fork of https://github.com/yanqingcheng/multi-war-atlas and implement it - just really make it super impressive and gorgeous and slick and performant and just absolutely NAIL it. deploy it to a qingsworkshop link when you're done. it's going to be a competition between models but please don't peek at the other guy's work

Am besten mit: Claude Opus 5.5 (Claude Code) und GPT-6 Sol (Codex) — jeweils identisch, parallel in T3 Code

Warum effektiv: Ein einziger locker formulierter Satz mit klaren Verben (fork, implement, deploy), expliziten Qualitätsanker (impressive, gorgeous, slick, performant) und einem Wettbewerbs-Kontext, der das Modell zu Höchstleistung animiert — inklusive Anti-Collusion-Klausel („don't peek at the other guy's work"). Die eigentliche Präzision steckt im Repo in einer INTENT.md-Brief-Datei: Der Prompt zeigt das aktuelle Muster „lockerer Kickoff-Satz + strukturierte Intent-Datei im Projekt".

Quelle: https://github.com/yanqingcheng/war-atlas-a-b | 5 GitHub-Sterne

Community Resonanz: Beide Agenten lieferten in einer Nacht interaktive Atlas-Websites mit vollständiger Commit-History — das Repo dokumentiert jedes Wort der menschlichen Nachrichten inklusive der Panik-Reaktion („oh jesus why would you not ask for the dataset") und ist damit ein ehrliches A/B-Experiment zum Lesen.

🖼️ TOP 3 PROMPTS — Bildgenerierung

1. „PE-T2I Photo Styles" — 17 fotografische Stile als Kamera-Sprache statt Fotografen-Name

Prompt (vollständig, kopierbar):

photo — a stray dog snaps its head around in a pitch-black alley after midnight rain
(Stil: Black Fury)

photo — a Tokyo street, hard evening light slicing a narrow alley, a passer-by pulling a long shadow, B&W high contrast
(Stil: Geometry of Light)

poster — a gritty B&W gig poster, huge distressed title "NOISE", a blurred running figure
(Stil: Black Fury, Poster-Modus)

poster — a B&W street-photography exhibition poster, bold title "STREET", a lone backlit figure swallowed by darkness
(Stil: Black Fury, Poster-Modus)

Am besten mit: Qwen-Image-2.1 (7B, lokal oder ComfyUI) — der PE-T2I-Rewriter expandiert die Kurzzeile zu vollem Prompt + Negative Prompt + Seitenverhältnis

Warum effektiv: Die Regel der Sammlung: Jeder Stil wird als „camera language + mood" beschrieben — nie über den Namen eines Fotografen. Subjekt, Text, Anzahl und Farben bleiben beim Nutzer, der Stil entscheidet nur über das, was man offengelassen hat; je kürzer der Prompt, desto stärker der Stil. Das ComfyUI-Node „PE-T2I Photo Style Prompt" schickt die Kurzzeile zusammen mit dem Style-Sheet an einen Rewriter (Qwen3.5-VL 9B, von Alibaba offiziell für genau diesen Zweck finetuned) und liefert langen Prompt, passenden Negative Prompt und die gewählte Bildgröße zurück.

Quelle: https://github.com/pottokao-dotcom/ComfyUI-QwenImage-PhotoStyles | 5 GitHub-Sterne

Community Resonanz: Die Galerie zeigt pro Stil zwei Fotos und zwei Poster — alles ungeachtet aus den Kurzzeilen generiert; das Repo betont „This is not a model loader": das Node formt nur Text und kostet weder VRAM noch Rechenzeit.

2. „Start-Still + Reveal-Still" — die zwei Referenzbilder, die jede Fly-Through-Video-Kette zusammenhalten

Prompt (vollständig, kopierbar):

Generate 2-3 start stills in different directions: an eye-level exterior view of the entrance with the doors open, inviting the camera in. Choose one (or let the user choose). This becomes the exact first frame.

Generate a reveal reference still from the chosen start still: a high aerial looking back at the rear of the same building, showing the exit opening, the yard, the full premises and its connected roads. This is the target for the final shot and keeps the architecture consistent.

Avoid real brand names, logos and distinctive trade dress in prompts (for example real car marques); request unbadged products and no signage lettering.

Am besten mit: GPT Image / Nano Banana 2.0 (via Higgsfield MCP) — dann als erstes/letztes Frame in die Videokette

Warum effektiv: Zwei Referenzbilder definieren Anfang und Ende der gesamten Kamerafahrt — das Start-Still wird exakter erster Frame, das Reveal-Still das Ziel des Schlussshots, und beide zusammen halten die Architektur über alle Clips konsistent. Die Anti-Brand-Regel („unbadged products, no signage lettering") verhindert Markenflecken in generierten Szenen, die später Content-Filter blockieren könnten.

Quelle: https://github.com/Barty-Bart/opus-55-10k-websites (prompts/01-desktop-drone-flythrough.md, Abschnitt „Stills first") | 4 GitHub-Sterne

Community Resonanz: Der Ansatz stammt aus dem gestern veröffentlichten Opus-5.5-Website-Workflow, in dem Stills zuerst vom Nutzer freigegeben werden, bevor teure Videoclips erzeugt werden — das Muster „Referenz-Still vor Video" übernehmen derzeit mehrere Pipeline-Skills.

3. „Handball-360°" — Foto zu frei drehbarem 3D-Rendering

Prompt (vollständig, kopierbar):

3D-render the handball court, goal, referee, players, and ball from the image so they can be viewed from any angle in a 360-degree view. Accurately reproduce each person's pose and the colors of the objects.

Am besten mit: Claude Opus 5.5 mit Bild-Input (oder GPT-6 Sol mit Vision), Ausgabe als interaktive 3D-Szene

Warum effektiv: Drei Sätze, drei präzise Aufträge: Was gerendert wird (Court, Goal, Schiedsrichter, Spieler, Ball), wie es benutzbar ist („viewed from any angle in a 360-degree view") und was die Qualität ausmacht („accurately reproduce each person's pose and the colors"). Kein Stil-Geschwafel — die Fidelity-Anforderung steht im Prompt, nicht im Auge des Betrachters. Die japanische Zweitfassung im selben Repo zeigt, dass der Prompt sprachunabhängig funktioniert.

Quelle: https://github.com/TripoGrowthLab/awesome-opus-5-5-prompts | 4 GitHub-Sterne

Community Resonanz: Der Prompt stammt aus derselben kuratierten Opus-5.5-Sammlung wie CatWalk — jede Kategorie (3D, Spiel, Animation, Simulation) ist dort mit source-verlinkten Beispielen belegt, damit Prompts nicht erfunden, sondern belegt werden.

🎬 TOP 3 PROMPTS — Videogenerierung

1. „/3dicon" — animiertes 3D-Icon als nahtloser Loop mit echtem Alpha

Prompt (vollständig, kopierbar):

make an animated 3d fire icon using /3dicon

Motion-Vorschlag des Skills (Format aus der Doku, wortgetreu):

**Motion** — the flame flickers and sways, throwing off tiny sparks
`event` · `lively` · `--emit`

Agree, or describe the motion you want?

Übersetzungsregeln für eigene Bewegungsbeschreibungen:

- what does this object do when left alone?  → --strategy
- how much should the object itself move?    → --energy
- would the action throw something off?     → --emit
- what literally happens to the material?   → --motion

Am besten mit: Claude Code mit installiertem 3dicon-Skill; Bild via GPT Image / Nano Banana, Motion via Seedance (OpenRouter), Ausgabe: animiertes WebP mit echtem weichem Alpha

Warum effektiv: Der Pipeline-Trick: dasselbe Still wird als erstes UND letztes Frame an Seedance geschickt, sodass das Video exakt dorthin zurückkehrt, wo es begann — der Loop schließt sich ohne sichtbare Naht. Die Kern-Disziplin steht im Skill-Prompt: „Always stop after the still" — erst ein Still (13 Cent), Freigabe abwarten, dann Motion (48 Cent, vier Minuten). Hintergrund wird gegen eine selbst gewählte Farbe keyed, damit die Originalfarben exakt gelöst statt geraten werden — weiche Kanten bleiben weich, ohne Halo.

Quelle: https://github.com/samyost1/3dicon | 272 GitHub-Sterne (seit zwei Tagen)

Community Resonanz: Einer der am schnellsten wachsenden Prompt-Skills der Woche; als Claude-Code-Plugin installierbar (/plugin marketplace add samyost1/3dicon) und direkt für App-Icons nutzbar, die expo-image, Chrome und Safari nativ rendern — ohne Lottie-Konvertierung.

2. „Nie das Wort drone schreiben" — die Kamera-Formel für lückenlose FPV-Flüge

Prompt (vollständig, kopierbar):

one single continuous first-person camera move, no cuts; the camera itself is flying; nothing flying or hovering is ever visible in frame.

every vehicle/person/object is stationary unless described; nothing appears, disappears or changes

no text, no logos, no brand badges, no signage lettering

Kette (Seedance 2.5, wortgetreu aus der Anleitung):

Clip A: image-to-video (Seedance 2.5 omni_reference) with the start still as start_image — entrance and main space
Clip B: video_extension, extension_mode: forward, with clip A as reference — back-of-house and storage
Clip C: extension of clip B plus reveal reference still — exit, 180° yaw, reverse climb

Reparatur statt Neugenerierung (FLUX 3 Video Edit):
remove the small flying quadcopter from every frame; keep everything else identical

Am besten mit: Seedance 2.5 (omni_reference + video_extension) via Higgsfield MCP, Repair via FLUX 3 Video Edit — gesteuert von Claude Opus 5.5

Warum effektiv: Die Anti-Drone-Regel ist die wichtigste Lektion des Tages: Schreibt man „drone", „quadcopter" oder „UAV" im Prompt, malt das Modell eine Drohne ins Bild — beschrieben werden muss die Kamera selbst. Dazu die Statik-Klausel (nichts erscheint oder verschwindet), getimete Manöver pro Segment und die Verkettung per Extension statt separater Szenen, weil Extension Kamera, Licht und Tempo vom letzten Frame des Vorgänger-Clips fortführt — genau das macht den Flug ununterbrochen. Reparaturen laufen gezielt („edit out small artefacts") statt als teure Endlos-Regenerierung.

Quelle: https://github.com/Barty-Bart/opus-55-10k-websites (prompts/01-desktop-drone-flythrough.md, Abschnitt „Prompt wording that matters") | 4 GitHub-Sterne

Community Resonanz: Der Workflow empfiehlt FFmpeg-Qualitätskontrolle pro Clip (Contact-Sheets, Scene-Cut-Detection select='gt(scene,0.3)') — die Reddit-ferne Prompt-Community nimmt diese Video-Disziplin aktuell breit auf.

3. „Cyberpunk-Megacity" — der Blender-Einzeiler mit kompletter Produktionspipeline

Prompt (vollständig, kopierbar):

build a complete cyberpunk megacity inside Blender with a hero train, procedural architecture, elevated rail systems, rain, volumetric atmosphere, cinematic lighting, multiple camera setups, and a full animated sequence.

Am besten mit: Claude Opus 5.5 (Agent mit Blender-Zugriff, z. B. über MCP oder als Skill)

Warum effektiv: Ein Satz, der wie eine Shot-List funktioniert: Erst das Asset (megacity, hero train), dann die Systeme (procedural architecture, elevated rail systems), dann die Atmosphäre (rain, volumetric atmosphere, cinematic lighting) und zuletzt die Auslieferung (multiple camera setups, full animated sequence). Jedes Komma ist eine Produktionsstufe — das Modell kann die Reihenfolge als Bau- und Renderplan lesen, ohne dass eine Zahl angepasst werden muss.

Quelle: https://github.com/TripoGrowthLab/awesome-opus-5-5-prompts | 4 GitHub-Sterne

Community Resonanz: Teil der gestern gestarteten Opus-5.5-Promptsammlung, in der 3D- und Animations-Prompts mit Quelle verlinkt sind — der Blender-Einzeiler wird dort neben variantenreicheren Mehrseiten-Prompts als Beispiel für maximale Wirkung pro Wort geführt.

🧠 TOP 3 NEUE TECHNIKEN

1. „Prompt-Fossils" — der Audit, der geschrieene CLAUDE.md-Regeln auf Normal-Lautstärke zurückdreht

Zusammenfassung: Ein Ein-Datei-Python-Tool findet die Zeilen in CLAUDE.md/AGENTS.md, die aktuelle Modelle wie Claude Opus 5.5 übererfüllen — und schreibt sie in normaler Lautstärke um.

Erklärung: Der Audit trennt „load-bearing"-Regeln von Fossilien: 388 der 460 Funde in den 1.000 most-starred GitHub-Repos sind sinnvolle Regeln, die nur in Versalien geschrien werden („NO EVIDENCE FILE == NO QA == NO COMMIT == NO PUSH. ALWAYS. EVERY TIME.") — und weil moderne Claude-Modelle Prompts extrem eng folgen, wird eine geschriene Regel überangewandt. Weitere Muster: nackte Verbotscluster („don't use emojis", 52 Funde), Druckverstärker („Be thorough - I am at risk of losing my job.", 14) und Step-Skripte für Urteilsarbeit. Der klassische Prompt-Hack ist ausgestorben: „think step by step" erscheint auf keiner der 35.105 Zeilen. --diff stellt die Lautstärke normal, ändert aber nichts am Inhalt; --max 0 macht den Audit CI-fähig.

Beispielprompt:

# Vorher (Fossil, echter Fund aus dem Korpus):
**NO EVIDENCE FILE == NO QA == NO COMMIT == NO PUSH.** ALWAYS. EVERY TIME.

# Nachher (Normalisierung per --diff — Regel bleibt, Schreien geht):
Run the evidence check before every commit and push.

Geeignet für: Alle aktuellen Agenten-Modelle (Claude Opus 5.5, GPT-6), CLAUDE.md/AGENTS.md und Skill-Dateien

Ursprung: https://github.com/stas4000/prompt-fossils

Warum heute wichtig: Mit Claude Opus 5.5 (seit Montag 40 % günstiger als Opus 5) folgen Agenten Regeln so eng, dass Prompt-Übersteuerung zum neuen Fehlerbild wird — der Audit über GitHubs Top-1.000 (23. September 2026) liefert erstmals Zahlen dazu: 339 der 1.000 Repos haben ein Root-Agent-File, 460 Funde in 101 Repos.

2. „Typed Decisions" — Prompts als typisierte Choice-Objekte statt Textgenerator

Zusammenfassung: Statt Freitext-Prompts definiert man benannte Felder mit Kriterien (Choice-Objekte) — das Modell antwortet mit Label, Level oder Wahrscheinlichkeitsverteilung, die Code komponieren kann.

Erklärung: Die Technik kommt aus den „System One"- bzw. Decision-Modellen (Jev von TypeSafe AI): Ein Prompt ist kein Fließtext, sondern ein Objekt mit instructions und criteria, wobei jedes Kriterium eine eigene Definition bekommt. Die Antwort trifft als getyptes Urteil mit kalibrierter Konfidenz ein („(count x peak - 1) / (count - 1)"), sodass Schwellwerte, Gewichte und Routing-Regeln in Python testbar bleiben — „a source of small, typed judgments that code composes, not as a text generator to be prompted and parsed". Gestern erschienen zwei Open-Alternativen: Contrastive-LMs CLM-8B (Agent-Aktionen bis zu 9× schneller bewertet als Jev) und Nokias AnyJev als training-freie Schicht, die jedes offene LLM in ein kalibriertes Decision-Modell verwandelt.

Beispielprompt:

INTENT = Choice(
    instructions="What the user wants the banking assistant to do",
    criteria={
        "check_balance": "See a balance or recent transactions",
        "approve_transfer": "Send or approve a transfer of money",
        "dispute_charge": "Contest a charge they do not recognise",
        "close_account": "Close the account permanently",
        "other": "Anything else, or not clear enough to act",
    },
)

Geeignet für: Jev (TypeSafe AI), CLM-8B (offene Gewichte), AnyJev (Layer auf jedem offenen LLM), awesome-jev listet 884 verifizierte Beispiele

Ursprung: https://www.marktechpost.com/2026/09/23/a-coding-guide-to-typesafe-ai-jev/ | https://github.com/Contrastive-LM/CLM

Warum heute wichtig: Mit CLM-8B und AnyJev ist die Technik seit gestern erstmals kostenlos und lokal machbar; der HN-Thread „Contrastive Language Models" (57 Punkte) und das 294-Sterne-CLM-Repo zeigen, dass Decision-Prompting vom Nischen- zum Overflow-Ansatz wird — vor allem für Agent-Routing und Support-Triage mit harten Budgetgrenzen.

3. „Scripted Vocal Bursts" — Bühnenregie mitten im TTS-Skript statt in separaten Beschreibungen

Zusammenfassung: Gemini 3.8 Flash TTS liest Regieanweisungen direkt aus dem Skript: nonverbale Cues in spitzen Klammern und Backchannel-Interjektionen in Pipes steuern Lachen, Seufzen und Zuhörer-Signale pro Zeile.

Erklärung: Statt Delivery in einem separaten Beschreibungsblock zu definieren, schreibt man die Regie direkt an die Stelle im Text, an der sie wirken soll: <laughs>, <sigh>, <gasp> für nichtverbale Ausbrüche, |mhm| oder |yeah| für aktives Zuhören — „for precise control" der Conversational-Textur. Dazu kommen Feinjustierungen der laufenden Performance per Natural Language („add subtle Southern US accent", „soften the delivery"), die zeilengenau greifen, sowie native Two-Speaker-Staging aus einem einzigen Skript. Damit wird TTS-Prompting vom „Voice-Presets auswählen" zum „Regisseur im Drehbuch".

Beispielprompt:

Add realistic conversational texture using non verbal cues (like <laughs>, <sigh>, <gasp>) and active-listening interjections (like |mhm| or |yeah|) for precise control.

Voice: a high-energy DJ voice from Melbourne
Delivery: soften the delivery, add subtle Southern US accent
Script: <laughs> You're tuned into the night shift, folks — |mhm| this next one's been on repeat all week. <gasp>

Geeignet für: Gemini 3.8 Flash TTS und Gemini 3.8 Flash-Lite TTS (AI Studio / Gemini API); das Muster übertragbar auf Qwen3-TTS und MOSS-TTS 2.0

Ursprung: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/

Warum heute wichtig: Google hat Voice-Design gestern von statischen Presets auf Natural-Language-Regie umgestellt (301 Punkte auf HN) — die Kombination aus Inline-Cues und zeilenweiser Delivery-Steuerung ist das TTS-Pendant zu Midjourneys Style-Referenzen und setzt sich als neuer Standard für Audiobooks, Podcasts und Voice Agents durch.

🏆 Highlight des Tages

„Voice Design per Prompt" — Gemini 3.8 Flash TTS macht Stimmen zu einer Natural-Language-Bauanleitung

Prompt (vollständig, kopierbar — Bausteine aus Googles Dokumentation und Demo-Galerie):

Voice: a high-energy DJ voice from Melbourne
Voice: a super-tinny, monotone robot voice
Voice: a dramatic, fire-breathing dragon with a distinct regional cadence (Japanese)

Remixing (Voice-Feintuning per Prompt):
"add subtle Southern US accent"
"soften the delivery"

Zwei-Sprecher-Szene aus einem Skript (Native Two-Speaker Scene Staging):
Speaker 1 [cheerful and friendly]: <laughs> Welcome back to the show!
Speaker 2 [calm and relaxed]: |mhm| Good to be here. <sigh> It's been a week.

Am besten mit: Gemini 3.8 Flash TTS (kreatives Voice-Design, 2.000+ Stimmen, 100+ Sprachen/Dialekte) und Gemini 3.8 Flash-Lite TTS (günstiges High-Volume-Dubbing)

Warum effektiv: Einmaliges Ziel, drei Hebel: Voices komplett neu aus Natürlicher Sprache erschaffen (Rolle, Akzent, Charakteristik), die laufende Delivery zeilengenau umformulieren („soften the delivery") und nonverbale Cues direkt ins Skript schreiben — alles in einem Prompt-Fenster statt in Preset-Menüs. Voice-Replication aus 30 Sekunden Audio kommt mit Consent-Verifikation, SynthID-Wasserzeichen und C2PA-Credentials. Simon Willison baute binnen Stunden einen Playground (BYO-API-Key) und generierte 1 m 18 s Zwei-Pelikaner-Dialog für 2,74 Cent.

Quelle: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ | 301 Upvotes, 137 Kommentare (https://news.ycombinator.com/item?id=49817615)

Community Resonanz: Die HN-Diskussion dreht sich um Preisdemokratisierung (250k-Wörter-Audiobuch ≈ 15 $ vs. 75 $ bei ElevenLabs) und ob die Beispiele wirklich tun, was der Prompt sagt — Stimmen-Schöpfer wie Audiobook-Nutzer testen gerade, wie präzise die Regie wirklich greift; Googles eigene Demo „ignoriert <chuckles>-Cues gelegentlich" ist der ehrlichste Fund des Threads.

📰 Erlesene Artikel & Ressourcen


Bericht erstellt am 24. September 2026 Quellen: Hacker News, AI News Portals, arXiv, GitHub, Personal Blogs