I found this project which is just a couple of small python scripts glueing various tools together: https://github.com/vndee/local-talking-llm

It’s pretty basic, but couldn’t find anything more polished. I did a little “vibe coding” to use a faster Chatterbox fork, stream the output back so I don’t have to wait for the entire LLM to finish before it starts “talking,” start recording on voice detection instead of the enter key, and allow interruption of the agent. But, like most vibe-coded stuff, it’s buggy. Was curious if there was something better that already exists before I commit to actually fixing the problems and pushing a fork.

  • hendrik@palaver.p3x.de
    link
    fedilink
    English
    arrow-up
    2
    ·
    6 days ago

    I got a bonus question… Is there a good end-to-end voice conversation solution? I’d like to try something which directly processes the audio and returns audio, rather than the whole pipeline with vad -> stt -> llm -> tts