Ollama: Build a ChatGPT-like Web Application

In this chapter, we will build a ChatGPT-like application: multi-session switching, history persistence to a database, a complete frontend interface, and three deployment methods.


Feature Planning and Data Model

A usable chat product needs at least four capabilities: multi-session isolation, history persistence, streaming replies, and session management (create and delete).

The database model supporting these capabilities only needs two tables:

多会话数据模型:chats 与 messages 一对多

chats stores the session list, and messages stores each message grouped by foreign key. When switching sessions, assemble the corresponding messages into Ollama's messages array to restore the full context.

The tech stack follows the minimal-dependency principle: Flask + sqlite3 (Python standard library) + vanilla JavaScript frontend, zero frontend frameworks.

Install flask and the ollama extension:

pip install flask ollama

Load the model:

ollama pull qwen3.5:4b

Backend: Complete API for Sessions and Messages

The backend has clear responsibilities: manage the two tables, and write each streaming reply from /api/chat to the database one by one.

Example

# File path: server.py
import json
import sqlite3
from flask import Flask, request, Response, jsonify, stream_with_context
from ollama import chat

app = Flask(__name__)
DB = 'chats.db'

def db():
    """Open a database connection; rows are returned as dictionaries"""
    conn = sqlite3.connect(DB)
    conn.row_factory = sqlite3.Row
    return conn

def init_db():
    with db() as conn:
        conn.executescript('''
        CREATE TABLE IF NOT EXISTS chats(
            id INTEGER PRIMARY KEY AUTOINCREMENT,
            title TEXT NOT NULL,
            created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
        );
        CREATE TABLE IF NOT EXISTS messages(
            id INTEGER PRIMARY KEY AUTOINCREMENT,
            chat_id INTEGER NOT NULL,
            role TEXT NOT NULL,
            content TEXT NOT NULL
        );
        '''
)

init_db()

# Session list
@app.get('/api/chats')
def list_chats():
    with db() as conn:
        rows = conn.execute(
            'SELECT id, title, created_at FROM chats ORDER BY id DESC'
        ).fetchall()
    return jsonify([dict(r) for r in rows])

# Create a new session
@app.post('/api/chats')
def create_chat():
    title = (request.json or {}).get('title', 'New session')
    with db() as conn:
        cur = conn.execute('INSERT INTO chats(title) VALUES (?)', (title,))
        return jsonify({'id': cur.lastrowid, 'title': title})

# Delete session (along with its messages)
@app.delete('/api/chats/<int:cid>')
def delete_chat(cid):
    with db() as conn:
        conn.execute('DELETE FROM messages WHERE chat_id=?', (cid,))
        conn.execute('DELETE FROM chats WHERE id=?', (cid,))
    return jsonify({'ok': True})

# Read the history of a session
@app.get('/api/chats/<int:cid>/messages')
def get_messages(cid):
    with db() as conn:
        rows = conn.execute(
            'SELECT role, content FROM messages WHERE chat_id=? ORDER BY id',
            (cid,)
        ).fetchall()
    return jsonify([dict(r) for r in rows])

Backend Core: Streaming Dialogue and History Write-back

The chat interface does three things: fetch history and build context, stream generation, and write the complete reply back to the database.

Example

# Append to server.py
@app.post('/api/chat')
def do_chat():
    data = request.get_json()
    cid, user_input = data['chat_id'], data['message']

    # Fetch history (including the just-inserted user message) as the model context
    with db() as conn:
        conn.execute(
            'INSERT INTO messages(chat_id, role, content) VALUES (?,?,?)',
            (cid, 'user', user_input)
        )
        history = [dict(r) for r in conn.execute(
            'SELECT role, content FROM messages WHERE chat_id=? ORDER BY id',
            (cid,)
        )]

    messages = [{'role': 'system',
                 'content': 'You are the EXAMPLE programming assistant. Answer accurately and concisely.'}]
    messages += history

    def generate():
        reply = ''
        stream = chat(model='qwen3.5:4b',
                      messages=messages, stream=True)
        for chunk in stream:
            reply += chunk.message.content
            # Push each chunk to the frontend as NDJSON
            yield json.dumps(
                {'delta': chunk.message.content},
                ensure_ascii=False) + '\n'
        # Key: write the complete reply back to the database to form persistent memory
        with db() as conn:
            conn.execute(
                'INSERT INTO messages(chat_id, role, content) VALUES (?,?,?)',
                (cid, 'assistant', reply))
        yield json.dumps({'done': True}, ensure_ascii=False) + '\n'

    return Response(
        stream_with_context(generate()),
        mimetype='application/x-ndjson')

# Homepage: return the frontend page
@app.get('/')
def index():
    return open('index.html', encoding='utf-8').read()

if __name__ == '__main__':
    app.run(port=5000)

Frontend: Chat Interface in Vanilla JS

The frontend is feature-complete and simple enough: session list on the left, message area plus input box on the right, with fetch streaming rendering.

Example

<!DOCTYPE html>
<!-- File path: index.html (same directory as server.py) -->
<html lang="zh-CN">
<head><meta charset="UTF-8"><title>Local ChatGPT</title>
<style>
  body { display:flex; height:100vh; margin:0; font-family: sans-serif; }
  aside { width:220px; border-right:1px solid #ddd; overflow-y:auto; }
  aside button { display:block; width:100%; text-align:left;
                 padding:10px; border:0; background:none; cursor:pointer; }
  aside button:hover { background:#f2f2f2; }
  main { flex:1; display:flex; flex-direction:column; }
  #log { flex:1; overflow-y:auto; padding:20px; }
  .msg { margin:8px 0; line-height:1.6; white-space:pre-wrap; }
  form { display:flex; border-top:1px solid #ddd; }
  input { flex:1; padding:12px; border:0; }
</style>
</head>
<body>
<aside>
  <button onclick="newChat()">+ New Session</button>
  <div id="list"></div>
</aside>
<main>
  <div id="log"></div>
  <form id="f">
    <input id="q" placeholder="Type a question, press Enter to send" autocomplete="off">
  </form>
</main>
<script>
let chatId = null;
const log = document.getElementById('log');
const list = document.getElementById('list');

// Load session list
async function loadChats() {
  const chats = await (await fetch('/api/chats')).json();
  list.innerHTML = '';
  for (const c of chats) {
    const b = document.createElement('button');
    b.textContent = c.title;
    b.onclick = () => openChat(c.id);
    list.appendChild(b);
  }
}

// Open a session and restore history
async function openChat(id) {
  chatId = id;
  log.innerHTML = '';
  const msgs = await (await fetch(`/api/chats/${id}/messages`)).json();
  for (const m of msgs) addMsg(m.role, m.content);
}

// Create a new session
async function newChat() {
  const c = await (await fetch('/api/chats', {method:'POST'})).json();
  await loadChats();
  openChat(c.id);
}

// Append a message to the interface
function addMsg(role, text) {
  const div = document.createElement('div');
  div.className = 'msg';
div.textContent = (role === 'user' ? 'You: ' : 'Assistant: ') + text;
  log.appendChild(div);
  log.scrollTop = log.scrollHeight;
  return div;
}

// Send and stream-render the reply
document.getElementById('f').onsubmit = async e => {
  e.preventDefault();
  const input = document.getElementById('q');
  const q = input.value.trim();
  if (!q || !chatId) return;
  input.value = '';
  addMsg('user', q);

  const div = addMsg('assistant', '');
  const resp = await fetch('/api/chat', {method:'POST',
    headers:{'Content-Type':'application/json'},
    body: JSON.stringify({chat_id: chatId, message: q})});
  const reader = resp.body.getReader();
  const dec = new TextDecoder();
  let buf = '';
  while (true) {
    const {done, value} = await reader.read();
    if (done) break;
    buf += dec.decode(value, {stream:true});
    const lines = buf.split('\n');
    buf = lines.pop();
    for (const line of lines) {
      if (!line) continue;
      const obj = JSON.parse(line);
      if (obj.delta) {
        div.textContent += obj.delta;
        log.scrollTop = log.scrollHeight;
      }
    }
  }
};

newChat();
loadChats();
</script>
</body>
</html>

Start and access:

python server.py

Open http://localhost:5000 in your browser to see the generated page.

Verify three core features: history is fully restored when switching sessions on the left; new sessions do not interfere with each other; history still exists after restarting server.py (persisted to chats.db).


Deployment: Three Tiers

TierApproachNotes
Local personal useRun python server.py directlyListens on 127.0.0.1 by default, safe
Intranet sharingapp.run(host='0.0.0.0'), members access http://server-IP:5000No authentication; be sure to add Nginx Basic Auth (see the private deployment chapter)
Long-term operation on cloud serversystemd service registration + Nginx reverse proxy + domain nameFirst run ollama pull <model> on the server

Use systemd on the cloud server to keep the application always running:

Example

# File path: /etc/systemd/system/example-chat.service
[Unit]
Description=Example Chat Web App
After=network.target ollama.service

[Service]
WorkingDirectory=/opt/example-chat
ExecStart=/usr/bin/python3 server.py
Restart=always
RestartSec=3

[Install]
WantedBy=multi-user.target

Set up long-term running:

sudo systemctl daemon-reload
sudo systemctl enable --now example-chat

The complete security checklist for external deployment (loopback port, reverse proxy authentication, exposure surface self-check) is in the Security and Compliance chapter; be sure to go through it before going live.

Other Extensions