본문 바로가기

Spring AI > 로컬 환경에서 Ollama를 이용하여 LLM 연동하기

@whateverU2026. 1. 20. 02:10
반응형

목적

로컬 데스크탑 환경에서 Ollama 를 통해 Qwen LLM 모델을 연동해 보자.

인터넷이 연결되어 있지 않은 환경에서도 LLM 모델을 로컬환경에서 돌려볼 수 있다.

 

Ollama 란? 

  • 모델 다운로드 관리자
    ollama pull qwen3-vl:8b
    ollama ps
  • 로컬 환경에서 LLM 실행을 담당하는 엔진 역할
  • LLM을 API 서버로 띄움
    curl http://localhost:11434/api/generate \
      -d '{"model":"qwen3-vl:8b","prompt":"안녕"}'

Ollama 설치

https://ollama.com/download

 

Download Ollama on macOS

Download Ollama for macOS

ollama.com

각 환경에 맞는 설치 파일을 다운로드 받아서 진행하면 된다.

LLM 모델 다운로드

llama4:16x17b 모델의 경우 67GB 이다.

16은 16개의 전문가 17b 는 17 Billion 

17B = 17 × 1,000,000,000 = 17,000,000,000 (170억) 개의 파라미터를 가진 모델이란 뜻이다.

https://ollama.com/library/llama4

 

아무 생각없이 해당 모델로 시도해보려고 했다가 아래와 같은 오류를 만났다.

500 Internal Server Error: model requires more system memory (59.5 GiB) than is available (6.6 GiB)
 
llama4 모델을 사용하기엔 Pc 스팩이 안되어서 가장 가성비 있는 모델인 llama3.1:8b 로 변경했다가 한국어 응답이 너무 부자연 스러워서 qwen3-vl:8b 모델로 최종 선택하였다.

 

qwen3-vl:8b 모델의 경우 멀티 모달에 대한 처리가 가능하며 (이미지 + 텍스트) 이미지에서 특정 요소를 인식하거나 텍스트 등을 추출 할 수 있다. 또한 영상도 이해가 가능하며 코드를 생성하는 기능까지 포함하고 있는 모델이다. 자세한건 해당 문서를 보면 알 수 있다. 

 

현재 우리나라에서 진행하고 있는 한국식 소버린 AI 콘테스트?를 보면 결국 해당 모델들을 만드는 작업이지 않을까 싶다 

로컬 환경에서 ollama 에서 spring 에서 기본 hello world 나오는 예제 알려줘 를 질문

 

Spring Ai Local LLM API 연동

build.gradle.kt

plugins {
    kotlin("jvm") version "1.9.25"
    kotlin("plugin.spring") version "1.9.25"
    id("org.springframework.boot") version "3.5.10-SNAPSHOT"
    id("io.spring.dependency-management") version "1.1.7"
}

group = "whatever.study.ai"
version = "0.0.1-SNAPSHOT"
description = "spring-ai"

java {
    toolchain {
        languageVersion = JavaLanguageVersion.of(17)
    }
}

repositories {
    mavenCentral()
    maven { url = uri("https://repo.spring.io/milestone") }
    maven { url = uri("https://repo.spring.io/snapshot") }
}

dependencies {
    implementation("org.springframework.boot:spring-boot-starter-web")
    implementation(platform("org.springframework.ai:spring-ai-bom:1.1.2"))
    implementation("org.springframework.ai:spring-ai-starter-mcp-server-webmvc")
    implementation("org.springframework.ai:spring-ai-starter-model-ollama")
    implementation("org.springframework.boot:spring-boot-starter-webflux")

    implementation("org.jetbrains.kotlin:kotlin-reflect")
    testImplementation("org.springframework.boot:spring-boot-starter-test")
    testImplementation("org.jetbrains.kotlin:kotlin-test-junit5")
    testRuntimeOnly("org.junit.platform:junit-platform-launcher")
}

kotlin {
    compilerOptions {
        freeCompilerArgs.addAll("-Xjsr305=strict")
    }
}

tasks.withType<Test> {
    useJUnitPlatform()
}

 

application.properties

spring.application.name=spring-ai
spring.ai.mcp.server.protocol=streamable

spring.ai.ollama.base-url=http://localhost:11434
spring.ai.ollama.chat.options.model=qwen3-vl:8b
spring.ai.ollama.chat.options.temperature=0.7

 

 

Controller

일반 ChatGpt 와 비슷하게 스트리밍 방식과 일반 동기식 요청으로 구현.

package whatever.study.ai.chat

import org.springframework.ai.chat.messages.UserMessage
import org.springframework.ai.chat.model.ChatModel
import org.springframework.ai.chat.prompt.Prompt
import org.springframework.http.MediaType
import org.springframework.web.bind.annotation.*
import org.springframework.web.servlet.mvc.method.annotation.SseEmitter
import java.io.IOException

data class ChatRequest(val message: String)
data class ChatResponse(val answer: String)

@RestController
@RequestMapping("/api/chat")
class ChatController(
    private val chatModel: ChatModel
) {
    @PostMapping
    fun chat(@RequestBody req: ChatRequest): ChatResponse {
        val answer = chatModel.call(req.message)
        return ChatResponse(answer)
    }

    @PostMapping(
        "/stream",
        produces = [MediaType.TEXT_EVENT_STREAM_VALUE],
        consumes = [MediaType.APPLICATION_JSON_VALUE]
    )
    fun stream(@RequestBody req: ChatRequest): SseEmitter {
        val emitter = SseEmitter(0L)

        val prompt = Prompt(listOf(UserMessage(req.message)))

        val disposable = chatModel.stream(prompt)
            .mapNotNull { it.result.output.text }
            .filter { it?.isNotBlank() == true }
            .subscribe(
                { chunk ->
                    try {
                        chunk?.let {
                            emitter.send(
                                SseEmitter.event()
                                    .name("token")
                                    .data(chunk)
                            )
                        }
                    } catch (e: IOException) {
                        emitter.completeWithError(e)
                    }
                },
                { err ->
                    try {
                        emitter.send(
                            SseEmitter.event()
                                .name("error")
                                .data(err.message ?: "unknown error")
                        )
                    } catch (_: Exception) {
                        // ignore
                    } finally {
                        emitter.completeWithError(err)
                    }
                },
                {
                    try {
                        emitter.send(
                            SseEmitter.event()
                                .name("done")
                                .data("[DONE]")
                        )
                    } catch (_: Exception) {
                        // ignore
                    } finally {
                        emitter.complete()
                    }
                }
            )

        emitter.onCompletion { disposable.dispose() }
        emitter.onTimeout {
            disposable.dispose()
            emitter.complete()
        }
        emitter.onError {
            disposable.dispose()
        }

        return emitter
    }
}

chat.html

<!doctype html>
<html lang="ko">
<head>
    <meta charset="utf-8" />
    <meta name="viewport" content="width=device-width, initial-scale=1" />
    <title>Ollama(Qwen) Chat Test</title>
    <style>
        body { font-family: system-ui, -apple-system, Segoe UI, Roboto, sans-serif; margin:0; background:#0b0f19; color:#e6e8ee; }
        .wrap { max-width: 900px; margin:0 auto; padding:24px; }
        .card { background:#11182a; border:1px solid #223055; border-radius:16px; padding:16px; box-shadow:0 8px 30px rgba(0,0,0,.25); }
        textarea { width:100%; min-height:90px; resize:vertical; padding:12px; border-radius:12px; border:1px solid #2b3b66; background:#0c1222; color:#e6e8ee; }
        button { padding:10px 14px; border-radius:12px; border:1px solid #2b3b66; background:#182446; color:#e6e8ee; cursor:pointer; }
        button:disabled { opacity:.6; cursor:not-allowed; }
        pre { white-space: pre-wrap; word-break: break-word; margin:0; font-family: ui-monospace, SFMono-Regular, Menlo, Monaco, Consolas, monospace; font-size:13px; line-height:1.5; }
        .row { display:flex; gap:12px; flex-wrap:wrap; margin-top:12px; }
        .hint { color:#aab3c7; font-size:12px; margin-top:10px; }
        .error { color:#ffb4b4; }
        code { background:#0c1222; padding:2px 6px; border-radius:8px; border:1px solid #2b3b66; }
    </style>
</head>
<body>
<div class="wrap">
    <h2 style="margin:0 0 14px 0;">Ollama(Qwen) Chat Test</h2>

    <div class="card" style="margin-bottom:14px;">
        <label for="msg" style="display:block; margin-bottom:8px;">질문</label>
        <textarea id="msg" placeholder="예) 한국어로 자연스럽게 대답해줘. Spring AI에서 Ollama 쓰는 법 알려줘"></textarea>

        <div class="row">
            <button id="btnOnce">일반 요청</button>
            <button id="btnStream">스트리밍 요청(SSE)</button>
            <button id="btnStop" disabled>스트리밍 중단</button>
            <button id="btnClear" type="button">지우기</button>
        </div>

        <div class="hint">
            엔드포인트: <code>POST /api/chat</code>, <code>POST /api/chat/stream</code>
            <span id="status" style="margin-left:10px;"></span>
        </div>
    </div>

    <div class="card">
        <div style="display:flex; justify-content:space-between; align-items:center; gap:12px; margin-bottom:10px;">
            <strong>답변</strong>
            <span class="hint">SSE는 <code>data:</code> 라인만 파싱해서 실시간으로 붙입니다.</span>
        </div>
        <pre id="out">답변이 여기에 표시됩니다.</pre>
        <div id="err" class="hint error" style="margin-top:10px;"></div>
    </div>
</div>

<script>
    const $ = (s) => document.querySelector(s);

    const out = $("#out");
    const err = $("#err");
    const status = $("#status");
    const msg = $("#msg");

    const btnOnce = $("#btnOnce");
    const btnStream = $("#btnStream");
    const btnStop = $("#btnStop");
    const btnClear = $("#btnClear");

    let abortController = null;

    function setBusy(busy) {
        btnOnce.disabled = busy;
        btnStream.disabled = busy;
        btnStop.disabled = !busy;
        status.textContent = busy ? "요청 중..." : "";
    }
    function setOutput(text) { out.textContent = text || ""; }
    function setError(text) { err.textContent = text || ""; }

    btnClear.addEventListener("click", () => {
        msg.value = "";
        setOutput("답변이 여기에 표시됩니다.");
        setError("");
    });

    btnStop.addEventListener("click", () => {
        if (abortController) abortController.abort();
    });

    // 일반 요청: POST /api/chat  { message } -> { answer }
    btnOnce.addEventListener("click", async () => {
        setError("");
        setOutput("");
        const message = msg.value.trim();
        if (!message) return setError("질문을 입력해 주세요.");

        try {
            setBusy(true);
            const res = await fetch("/api/chat", {
                method: "POST",
                headers: { "Content-Type": "application/json" },
                body: JSON.stringify({ message })
            });

            const ct = res.headers.get("content-type") || "";
            if (!res.ok) {
                const body = await res.text().catch(() => "");
                throw new Error(`HTTP ${res.status} ${res.statusText}\n${body}`);
            }

            if (ct.includes("application/json")) {
                const data = await res.json();
                setOutput(data.answer ?? JSON.stringify(data, null, 2));
            } else {
                setOutput(await res.text());
            }
        } catch (e) {
            setError(String(e));
        } finally {
            setBusy(false);
        }
    });

    // 스트리밍 요청: POST /api/chat/stream (TEXT_EVENT_STREAM)
    // SSE는 "data: ...\n\n" 형태라 data 라인만 파싱해서 붙임
    btnStream.addEventListener("click", async () => {
        setError("");
        setOutput("");
        const message = msg.value.trim();
        if (!message) return setError("질문을 입력해 주세요.");

        abortController = new AbortController();

        try {
            setBusy(true);

            const res = await fetch("/api/chat/stream", {
                method: "POST",
                headers: { "Content-Type": "application/json" },
                body: JSON.stringify({ message }),
                signal: abortController.signal
            });

            if (!res.ok) {
                const body = await res.text().catch(() => "");
                throw new Error(`HTTP ${res.status} ${res.statusText}\n${body}`);
            }
            if (!res.body) throw new Error("ReadableStream을 사용할 수 없습니다. 브라우저를 확인해 주세요.");

            const reader = res.body.getReader();
            const decoder = new TextDecoder("utf-8");

            let buffer = "";
            let assembled = "";

            while (true) {
                const { done, value } = await reader.read();
                if (done) break;

                buffer += decoder.decode(value, { stream: true });

                // 이벤트 구분은 \n\n
                const events = buffer.split("\n\n");
                buffer = events.pop() ?? "";

                for (const evt of events) {
                    const dataLines = evt
                        .split("\n")
                        .filter(line => line.startsWith("data:"))
                        .map(line => line.replace(/^data:\s?/, ""));

                    if (dataLines.length === 0) continue;

                    assembled += dataLines.join("\n");
                    setOutput(assembled);
                }
            }

            // SSE가 아닌 일반 텍스트 스트림이면 남은 버퍼 표시
            if (!assembled && buffer.trim()) setOutput(buffer);

        } catch (e) {
            if (String(e).includes("AbortError")) {
                setError("스트리밍을 중단했습니다.");
            } else {
                setError(String(e));
            }
        } finally {
            abortController = null;
            setBusy(false);
        }
    });
</script>
</body>
</html>

 

LLM 호출 테스트

대략 1분정도 소요됨..ㅎㅎㅎ (1분부터 보세요)

동영상 서비스가 종료되어 해당 콘텐츠를 재생할 수 없습니다.

 

 

결론

인터넷 없이도 쓸만한 LLM을 로컬환경에서 돌릴 수 있다. 

ChatGPT, Claude, Gemini 같은 서비스도 결국 LLM + 주변 시스템의 조합이지 않을까란 생각이 들었다.

 

LLM 모델 생성 방식이나 동작 과정등도 공부하면 좋을 것 같다.

참고문서

spring ai > ollama 연동 문서 https://docs.spring.io/spring-ai/reference/api/chat/ollama-chat.html

 

Ollama Chat :: Spring AI Reference

Ollama supports thinking mode for reasoning models that can emit their internal reasoning process before providing a final answer. This feature is available for models like Qwen3, DeepSeek-v3.1, DeepSeek R1, and GPT-OSS. Thinking mode helps you understand

docs.spring.io

ollama > open model download https://ollama.com/download/windows

 

Download Ollama on Windows

Download Ollama for Windows

ollama.com

 

 

반응형
whateverU
@whateverU :: whateverU

sang12.co.kr https://github.com/ChoiSangIl

공감하셨다면 ❤️ 구독도 환영합니다! 🤗

목차