FUTURE WORK
Beyond activity evidence
Context-aware decisions for Conversation Flow.
MicroCF provides a working basis for acoustic turn-taking. Our next direction, MacroCF, explores how an autoregressive audio–text model could use conversational context to decide when to respond, wait, acknowledge, or abstain.
- MacroCF: model direction
Accumulated audio and text context, a target decision cadence of 200 ms, and optional reasoning output. - Event interface
Five proposed control tags: EOT, INT, LTN, BC, and REJ, with optional <think> output.
Role of MicroCF
MicroCF is intended to remain the compact path for baseline services. It can also propose candidate events and ambiguous segments for annotation. Such candidates require human review and quality control before use as training labels; held-out evaluation data must remain separate.