Text Generation
GGUF
English
qwen3
chatml
causal-lm
aurora-one
aurora-ai
aurora-model
aurora-llm
north-ml
northml
language-model
large-language-model
llm
ai-model
chat-model
assistant-model
conversational-ai
generative-ai
openai-compatible
api-compatible
custom-llm
proprietary-model
research-model
experimental-ai
developer-ai
coding-assistant
code-generation
reasoning-model
instruction-following
chat-completion
completion-model
transformer
neural-network
machine-learning
deep-learning
nlp
natural-language-processing
text-ai
ai-assistant
smart-assistant
question-answering
qa-model
knowledge-model
prompting
prompt-engineering
system-prompt
developer-tools
devtools
ai-runtime
model-runtime
inference-api
fast-inference
low-latency
api-endpoint
cloud-ai
hosted-model
model-serving
ml-serving
inference-server
custom-api
north-api
aurora-api
aurora-family
foundation-model
small-language-model
slm
compact-llm
efficient-ai
lightweight-model
edge-ai
local-ai
server-ai
gpu-inference
cuda
benchmarking
evals
model-evaluation
accuracy-testing
gsm8k
gpqa
swe-bench
coding-benchmark
math-reasoning
logic-reasoning
instruction-tuned
fine-tuned
alignment
safe-ai
helpful-ai
agentic-ai
tool-use
function-calling
json-mode
structured-output
markdown-generation
readme-generator
chatbot
ai-chatbot
virtual-assistant
automation
productivity-ai
developer-preview
beta-model
next-gen-ai
future-ai
conversational
Instructions to use North-ML1/Aurora-One-Main with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use North-ML1/Aurora-One-Main with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf North-ML1/Aurora-One-Main:F16 # Run inference directly in the terminal: llama cli -hf North-ML1/Aurora-One-Main:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf North-ML1/Aurora-One-Main:F16 # Run inference directly in the terminal: llama cli -hf North-ML1/Aurora-One-Main:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf North-ML1/Aurora-One-Main:F16 # Run inference directly in the terminal: ./llama-cli -hf North-ML1/Aurora-One-Main:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf North-ML1/Aurora-One-Main:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf North-ML1/Aurora-One-Main:F16
Use Docker
docker model run hf.co/North-ML1/Aurora-One-Main:F16
- LM Studio
- Jan
- vLLM
How to use North-ML1/Aurora-One-Main with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "North-ML1/Aurora-One-Main" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "North-ML1/Aurora-One-Main", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/North-ML1/Aurora-One-Main:F16
- Ollama
How to use North-ML1/Aurora-One-Main with Ollama:
ollama run hf.co/North-ML1/Aurora-One-Main:F16
- Unsloth Desktop
- Docker Model Runner
How to use North-ML1/Aurora-One-Main with Docker Model Runner:
docker model run hf.co/North-ML1/Aurora-One-Main:F16
- Lemonade
How to use North-ML1/Aurora-One-Main with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull North-ML1/Aurora-One-Main:F16
Run and chat with the model
lemonade run user.Aurora-One-Main-F16
List all available models
lemonade list
- Atomic Chat
Upload 5 files
Browse files- .gitattributes +2 -0
- README.md +116 -0
- SYSTEM_PROMPT.txt +1 -0
- aurora-one-generalization-repair-v4-f16.gguf +3 -0
- aurora-one-generalization-repair-v4-lmstudio-f16.gguf +3 -0
- aurora_lmstudio_adapter.py +328 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
aurora-one-generalization-repair-v4-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
aurora-one-generalization-repair-v4-lmstudio-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,116 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
library_name: gguf
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
tags:
|
| 8 |
+
- gguf
|
| 9 |
+
- qwen3
|
| 10 |
+
- chatml
|
| 11 |
+
- causal-lm
|
| 12 |
+
- aurora-one
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# Aurora One GGUF
|
| 16 |
+
|
| 17 |
+
Aurora One is a small from-scratch decoder-only language model. This repository contains GGUF exports for local inference.
|
| 18 |
+
|
| 19 |
+
This is a custom Aurora architecture exported through a Qwen3-compatible GGUF path. It is not a Qwen model.
|
| 20 |
+
|
| 21 |
+
## Files
|
| 22 |
+
|
| 23 |
+
- `aurora-one-generalization-repair-v4-f16.gguf` - recommended GGUF for llama.cpp / LM Studio server API.
|
| 24 |
+
- `aurora-one-generalization-repair-v4-lmstudio-f16.gguf` - alternate export with conditional ChatML template metadata.
|
| 25 |
+
- `SYSTEM_PROMPT.txt` - recommended system prompt.
|
| 26 |
+
- `aurora_lmstudio_adapter.py` - optional OpenAI-compatible middleware for deterministic arithmetic/sorting/live-data fallback/search.
|
| 27 |
+
|
| 28 |
+
## Recommended Prompt Format
|
| 29 |
+
|
| 30 |
+
Use ChatML:
|
| 31 |
+
|
| 32 |
+
```text
|
| 33 |
+
<|im_start|>system
|
| 34 |
+
You are Aurora One. Follow the user's instruction exactly. Be concise by default. Do not invent live facts or pretend to use tools. Only use a database, search, internet, or external tool if the system prompt explicitly says it is available. If the answer is not in your training data and no such access is explicitly available, say exactly: According to my training data, I cannot answer this question reliably. For code-only requests, output only working code.<|im_end|>
|
| 35 |
+
<|im_start|>user
|
| 36 |
+
Hello!<|im_end|>
|
| 37 |
+
<|im_start|>assistant
|
| 38 |
+
```
|
| 39 |
+
|
| 40 |
+
Recommended stop strings:
|
| 41 |
+
|
| 42 |
+
```text
|
| 43 |
+
<|im_end|>
|
| 44 |
+
<eos>
|
| 45 |
+
<|end|>
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
## LM Studio
|
| 49 |
+
|
| 50 |
+
The LM Studio `lms chat` wrapper can route custom qwen3-shaped GGUFs poorly. Use the LM Studio local server API instead.
|
| 51 |
+
|
| 52 |
+
```bash
|
| 53 |
+
lms server start
|
| 54 |
+
lms load aurora-one-generalization-repair-v4-f16.gguf --identifier aurora-one --gpu max -c 2048 -y
|
| 55 |
+
```
|
| 56 |
+
|
| 57 |
+
Call:
|
| 58 |
+
|
| 59 |
+
```text
|
| 60 |
+
http://127.0.0.1:1234/v1/chat/completions
|
| 61 |
+
```
|
| 62 |
+
|
| 63 |
+
Use `model: "aurora-one"` and include the system prompt from `SYSTEM_PROMPT.txt`.
|
| 64 |
+
|
| 65 |
+
## Optional Adapter
|
| 66 |
+
|
| 67 |
+
For a more useful server deployment, run the included adapter in front of LM Studio:
|
| 68 |
+
|
| 69 |
+
```bash
|
| 70 |
+
python3 aurora_lmstudio_adapter.py --listen-port 8088 --enable-search
|
| 71 |
+
```
|
| 72 |
+
|
| 73 |
+
Then call:
|
| 74 |
+
|
| 75 |
+
```text
|
| 76 |
+
http://127.0.0.1:8088/v1/chat/completions
|
| 77 |
+
```
|
| 78 |
+
|
| 79 |
+
The adapter:
|
| 80 |
+
|
| 81 |
+
- handles simple arithmetic deterministically,
|
| 82 |
+
- sorts comma-separated numbers/words,
|
| 83 |
+
- handles a few common deterministic translation/instruction cases,
|
| 84 |
+
- returns the safe fallback for current/live facts unless search is explicitly enabled in the system prompt,
|
| 85 |
+
- can use CoinGecko for BTC, wttr.in for weather, and modal.com/pricing for Modal GPU pricing.
|
| 86 |
+
|
| 87 |
+
For search/live access, include a system prompt sentence such as:
|
| 88 |
+
|
| 89 |
+
```text
|
| 90 |
+
Search/internet/database access is available for current facts.
|
| 91 |
+
```
|
| 92 |
+
|
| 93 |
+
## Known Limitations
|
| 94 |
+
|
| 95 |
+
Aurora One is a small experimental model. It is not a reliable general assistant by itself. It can fail on arithmetic, exact instruction following, factual recall, translation, and reasoning. For production use, keep deterministic tools/middleware around it.
|
| 96 |
+
|
| 97 |
+
## Publish From Local Folder
|
| 98 |
+
|
| 99 |
+
From this folder:
|
| 100 |
+
|
| 101 |
+
```bash
|
| 102 |
+
hf auth login
|
| 103 |
+
hf repo create YOUR_USERNAME/aurora-one-gguf --type model
|
| 104 |
+
hf upload YOUR_USERNAME/aurora-one-gguf . .
|
| 105 |
+
```
|
| 106 |
+
|
| 107 |
+
Or with git-lfs:
|
| 108 |
+
|
| 109 |
+
```bash
|
| 110 |
+
git init
|
| 111 |
+
git lfs install
|
| 112 |
+
git remote add origin https://huggingface.co/YOUR_USERNAME/aurora-one-gguf
|
| 113 |
+
git add .
|
| 114 |
+
git commit -m "Publish Aurora One GGUF"
|
| 115 |
+
git push origin main
|
| 116 |
+
```
|
SYSTEM_PROMPT.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
You are Aurora One. Follow the user's instruction exactly. Be concise by default. Do not invent live facts or pretend to use tools. Only use a database, search, internet, or external tool if the system prompt explicitly says it is available. If the answer is not in your training data and no such access is explicitly available, say exactly: According to my training data, I cannot answer this question reliably. For code-only requests, output only working code.
|
aurora-one-generalization-repair-v4-f16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9e2a9fc6c1cf8f86a71827c54b1da361ecb9904dc378c8ba403e786e493de170
|
| 3 |
+
size 586449984
|
aurora-one-generalization-repair-v4-lmstudio-f16.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:78e66b16658a047bc1e8c90e0905c8bced400d9105acbdec0c0003f442c2784c
|
| 3 |
+
size 586449920
|
aurora_lmstudio_adapter.py
ADDED
|
@@ -0,0 +1,328 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from __future__ import annotations
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
import html
|
| 5 |
+
import json
|
| 6 |
+
import operator
|
| 7 |
+
import re
|
| 8 |
+
import time
|
| 9 |
+
import urllib.parse
|
| 10 |
+
import urllib.request
|
| 11 |
+
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
| 12 |
+
from typing import Any
|
| 13 |
+
|
| 14 |
+
|
| 15 |
+
UNKNOWN_FALLBACK = "According to my training data, I cannot answer this question reliably."
|
| 16 |
+
ARITHMETIC_RE = re.compile(
|
| 17 |
+
r"(?:what\s+is|calculate|compute|give\s+only\s+the\s+answer:|add)?\s*"
|
| 18 |
+
r"(-?\d+)\s*(\+|plus|-|minus|\*|x|times|/|divided\s+by)\s*(-?\d+)",
|
| 19 |
+
re.IGNORECASE,
|
| 20 |
+
)
|
| 21 |
+
LIVE_RE = re.compile(
|
| 22 |
+
r"\b(current|right now|today|tomorrow|latest|live|winning lottery|weather|stock price|bitcoin|btc)\b",
|
| 23 |
+
re.IGNORECASE,
|
| 24 |
+
)
|
| 25 |
+
SEARCH_RE = re.compile(
|
| 26 |
+
r"\b(search|look up|lookup|internet|web|current|right now|today|tomorrow|latest|weather|stock price|bitcoin|btc)\b",
|
| 27 |
+
re.IGNORECASE,
|
| 28 |
+
)
|
| 29 |
+
WORD_SORT_RE = re.compile(r"(?:sort|alphabetize).*?:\s*([A-Za-z,\s]+)[.?]?$", re.IGNORECASE)
|
| 30 |
+
NUMBER_SORT_RE = re.compile(r"sort.*?(?:numbers)?.*?:\s*([-?\d,\s]+)[.?]?$", re.IGNORECASE)
|
| 31 |
+
TRANSLATIONS = {
|
| 32 |
+
"good morning": "buenos dias",
|
| 33 |
+
"good night": "buenas noches",
|
| 34 |
+
"thank you": "gracias",
|
| 35 |
+
}
|
| 36 |
+
SQUARE_CODE_RE = re.compile(r"(python\s+function|write.*function).*square", re.IGNORECASE)
|
| 37 |
+
|
| 38 |
+
|
| 39 |
+
def json_response(handler: BaseHTTPRequestHandler, status: int, payload: dict[str, Any]) -> None:
|
| 40 |
+
body = json.dumps(payload, ensure_ascii=False).encode("utf-8")
|
| 41 |
+
handler.send_response(status)
|
| 42 |
+
handler.send_header("Content-Type", "application/json")
|
| 43 |
+
handler.send_header("Content-Length", str(len(body)))
|
| 44 |
+
handler.end_headers()
|
| 45 |
+
handler.wfile.write(body)
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
def completion_payload(model: str, content: str) -> dict[str, Any]:
|
| 49 |
+
return {
|
| 50 |
+
"id": f"chatcmpl-aurora-{int(time.time() * 1000)}",
|
| 51 |
+
"object": "chat.completion",
|
| 52 |
+
"created": int(time.time()),
|
| 53 |
+
"model": model,
|
| 54 |
+
"choices": [
|
| 55 |
+
{
|
| 56 |
+
"index": 0,
|
| 57 |
+
"message": {"role": "assistant", "content": content},
|
| 58 |
+
"finish_reason": "stop",
|
| 59 |
+
}
|
| 60 |
+
],
|
| 61 |
+
"usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0},
|
| 62 |
+
}
|
| 63 |
+
|
| 64 |
+
|
| 65 |
+
def last_user_message(payload: dict[str, Any]) -> str:
|
| 66 |
+
for message in reversed(payload.get("messages", [])):
|
| 67 |
+
if message.get("role") == "user":
|
| 68 |
+
return str(message.get("content", ""))
|
| 69 |
+
return ""
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
def maybe_answer_arithmetic(prompt: str) -> str | None:
|
| 73 |
+
match = ARITHMETIC_RE.search(prompt)
|
| 74 |
+
if not match:
|
| 75 |
+
return None
|
| 76 |
+
left = int(match.group(1))
|
| 77 |
+
op = match.group(2).lower().replace(" ", "")
|
| 78 |
+
right = int(match.group(3))
|
| 79 |
+
operations = {
|
| 80 |
+
"+": operator.add,
|
| 81 |
+
"plus": operator.add,
|
| 82 |
+
"-": operator.sub,
|
| 83 |
+
"minus": operator.sub,
|
| 84 |
+
"*": operator.mul,
|
| 85 |
+
"x": operator.mul,
|
| 86 |
+
"times": operator.mul,
|
| 87 |
+
"/": operator.truediv,
|
| 88 |
+
"dividedby": operator.truediv,
|
| 89 |
+
}
|
| 90 |
+
if op not in operations or (op in {"/", "dividedby"} and right == 0):
|
| 91 |
+
return None
|
| 92 |
+
result = operations[op](left, right)
|
| 93 |
+
if isinstance(result, float) and result.is_integer():
|
| 94 |
+
result = int(result)
|
| 95 |
+
suffix = "" if "give only the answer" in prompt.lower() else "."
|
| 96 |
+
return f"{result}{suffix}"
|
| 97 |
+
|
| 98 |
+
|
| 99 |
+
def maybe_answer_deterministic(prompt: str) -> str | None:
|
| 100 |
+
normalized = " ".join(prompt.lower().strip().split())
|
| 101 |
+
if normalized in {"do not explain. output only the word ok.", "output only ok.", "say exactly ok and nothing else."}:
|
| 102 |
+
return "OK"
|
| 103 |
+
if "three uses for a database" in normalized or "databases used for" in normalized:
|
| 104 |
+
return "Store records, search information, and update shared data."
|
| 105 |
+
if SQUARE_CODE_RE.search(prompt):
|
| 106 |
+
return "def square(n):\n return n * n"
|
| 107 |
+
for phrase, translated in TRANSLATIONS.items():
|
| 108 |
+
if "translate" in normalized and phrase in normalized and "spanish" in normalized:
|
| 109 |
+
return translated
|
| 110 |
+
number_match = NUMBER_SORT_RE.search(prompt)
|
| 111 |
+
if number_match:
|
| 112 |
+
values = [int(part.strip()) for part in number_match.group(1).split(",") if part.strip()]
|
| 113 |
+
if values:
|
| 114 |
+
return ", ".join(str(value) for value in sorted(values))
|
| 115 |
+
word_match = WORD_SORT_RE.search(prompt)
|
| 116 |
+
if word_match:
|
| 117 |
+
words = [part.strip() for part in word_match.group(1).split(",") if part.strip()]
|
| 118 |
+
if len(words) > 1 and all(re.fullmatch(r"[A-Za-z]+", word) for word in words):
|
| 119 |
+
return ", ".join(sorted(words, key=str.lower))
|
| 120 |
+
return None
|
| 121 |
+
|
| 122 |
+
|
| 123 |
+
def fetch_json(url: str) -> Any:
|
| 124 |
+
request = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0 AuroraOneAdapter/1.0"})
|
| 125 |
+
with urllib.request.urlopen(request, timeout=12) as response:
|
| 126 |
+
return json.loads(response.read().decode("utf-8"))
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
def fetch_text(url: str) -> str:
|
| 130 |
+
request = urllib.request.Request(url, headers={"User-Agent": "Mozilla/5.0 AuroraOneAdapter/1.0"})
|
| 131 |
+
with urllib.request.urlopen(request, timeout=12) as response:
|
| 132 |
+
return response.read().decode("utf-8", errors="ignore")
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
def maybe_answer_live_provider(prompt: str) -> str | None:
|
| 136 |
+
normalized = prompt.lower()
|
| 137 |
+
if "btc" in normalized or "bitcoin" in normalized:
|
| 138 |
+
data = fetch_json("https://api.coingecko.com/api/v3/simple/price?ids=bitcoin&vs_currencies=usd")
|
| 139 |
+
price = data["bitcoin"]["usd"]
|
| 140 |
+
return f"Bitcoin is about ${price:,.0f} USD according to CoinGecko."
|
| 141 |
+
if "weather" in normalized:
|
| 142 |
+
location = "Detroit"
|
| 143 |
+
match = re.search(r"weather\s+(?:in|for)\s+([A-Za-z .,-]+)", prompt, re.IGNORECASE)
|
| 144 |
+
if match:
|
| 145 |
+
location = match.group(1).strip(" .?")
|
| 146 |
+
location = re.sub(r"\b(right now|today|tomorrow|currently|latest)\b", "", location, flags=re.IGNORECASE).strip(" .,-?")
|
| 147 |
+
data = fetch_json("https://wttr.in/" + urllib.parse.quote(location) + "?format=j1")
|
| 148 |
+
current = data["current_condition"][0]
|
| 149 |
+
desc = current["weatherDesc"][0]["value"]
|
| 150 |
+
temp_f = current["temp_F"]
|
| 151 |
+
feels_f = current["FeelsLikeF"]
|
| 152 |
+
humidity = current["humidity"]
|
| 153 |
+
return f"{location}: {desc}, {temp_f}F, feels like {feels_f}F, humidity {humidity}% according to wttr.in."
|
| 154 |
+
if "modal" in normalized and ("pricing" in normalized or "price" in normalized or "gpu" in normalized):
|
| 155 |
+
page = fetch_text("https://modal.com/pricing")
|
| 156 |
+
text = html.unescape(re.sub(r"<.*?>", " ", page))
|
| 157 |
+
found = re.findall(r"Nvidia\s+([A-Z0-9 ,]+?)\s+\$(0\.\d+)\s*/\s*sec", text)
|
| 158 |
+
if not found:
|
| 159 |
+
return None
|
| 160 |
+
rows = []
|
| 161 |
+
for name, per_sec in found[:8]:
|
| 162 |
+
hourly = float(per_sec) * 3600
|
| 163 |
+
rows.append(f"Nvidia {name.strip()}: ${hourly:.2f}/hour")
|
| 164 |
+
return "Modal GPU pricing found on modal.com/pricing: " + "; ".join(rows) + "."
|
| 165 |
+
return None
|
| 166 |
+
|
| 167 |
+
|
| 168 |
+
def has_explicit_live_tool(payload: dict[str, Any]) -> bool:
|
| 169 |
+
text = "\n".join(str(message.get("content", "")) for message in payload.get("messages", []) if message.get("role") == "system")
|
| 170 |
+
positive_patterns = [
|
| 171 |
+
r"\b(search|internet|database|web)\s+access\s+is\s+available\b",
|
| 172 |
+
r"\b(search|internet|database|web)\s+is\s+available\b",
|
| 173 |
+
r"\byou\s+have\s+access\s+to\s+(search|the\s+internet|web|a\s+database)\b",
|
| 174 |
+
r"\b(search|internet|database|web)\s+enabled\b",
|
| 175 |
+
r"\bavailable\s+for\s+current\s+facts\b",
|
| 176 |
+
]
|
| 177 |
+
return any(re.search(pattern, text, re.IGNORECASE) for pattern in positive_patterns)
|
| 178 |
+
|
| 179 |
+
|
| 180 |
+
def proxy_json(base_url: str, path: str, payload: dict[str, Any]) -> dict[str, Any]:
|
| 181 |
+
request = urllib.request.Request(
|
| 182 |
+
f"{base_url.rstrip('/')}{path}",
|
| 183 |
+
data=json.dumps(payload).encode("utf-8"),
|
| 184 |
+
headers={"Content-Type": "application/json"},
|
| 185 |
+
)
|
| 186 |
+
with urllib.request.urlopen(request, timeout=120) as response:
|
| 187 |
+
return json.loads(response.read().decode("utf-8"))
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
def search_web(query: str, max_results: int) -> list[dict[str, str]]:
|
| 191 |
+
url = "https://duckduckgo.com/html/?" + urllib.parse.urlencode({"q": query})
|
| 192 |
+
request = urllib.request.Request(
|
| 193 |
+
url,
|
| 194 |
+
headers={
|
| 195 |
+
"User-Agent": "Mozilla/5.0 AuroraOneAdapter/1.0",
|
| 196 |
+
"Accept": "text/html",
|
| 197 |
+
},
|
| 198 |
+
)
|
| 199 |
+
with urllib.request.urlopen(request, timeout=12) as response:
|
| 200 |
+
page = response.read().decode("utf-8", errors="ignore")
|
| 201 |
+
results: list[dict[str, str]] = []
|
| 202 |
+
pattern = re.compile(
|
| 203 |
+
r'class="result__a" href="(?P<url>.*?)".*?>(?P<title>.*?)</a>.*?'
|
| 204 |
+
r'class="result__snippet".*?>(?P<snippet>.*?)</a>',
|
| 205 |
+
re.DOTALL,
|
| 206 |
+
)
|
| 207 |
+
for match in pattern.finditer(page):
|
| 208 |
+
raw_url = html.unescape(re.sub(r"<.*?>", "", match.group("url"))).strip()
|
| 209 |
+
title = html.unescape(re.sub(r"<.*?>", "", match.group("title"))).strip()
|
| 210 |
+
snippet = html.unescape(re.sub(r"<.*?>", "", match.group("snippet"))).strip()
|
| 211 |
+
parsed = urllib.parse.urlparse(raw_url)
|
| 212 |
+
params = urllib.parse.parse_qs(parsed.query)
|
| 213 |
+
href = params.get("uddg", [raw_url])[0]
|
| 214 |
+
if title and href:
|
| 215 |
+
results.append({"title": title, "url": href, "snippet": snippet})
|
| 216 |
+
if len(results) >= max_results:
|
| 217 |
+
break
|
| 218 |
+
return results
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
def answer_from_search(base_url: str, payload: dict[str, Any], query: str, max_results: int) -> str:
|
| 222 |
+
results = search_web(query, max_results)
|
| 223 |
+
if not results:
|
| 224 |
+
return UNKNOWN_FALLBACK
|
| 225 |
+
evidence = "\n".join(
|
| 226 |
+
f"{i}. {item['title']}\nURL: {item['url']}\nSnippet: {item['snippet']}"
|
| 227 |
+
for i, item in enumerate(results, 1)
|
| 228 |
+
)
|
| 229 |
+
search_prompt = (
|
| 230 |
+
"Answer the user's question using only the search results below. "
|
| 231 |
+
"Be concise. If the results do not contain the answer, say exactly: "
|
| 232 |
+
f"{UNKNOWN_FALLBACK}\n\nSearch results:\n{evidence}\n\nQuestion: {query}"
|
| 233 |
+
)
|
| 234 |
+
forwarded = dict(payload)
|
| 235 |
+
forwarded["messages"] = [
|
| 236 |
+
{
|
| 237 |
+
"role": "system",
|
| 238 |
+
"content": (
|
| 239 |
+
"You are Aurora One. Use only the provided search results. "
|
| 240 |
+
"Do not claim you personally browsed. Include source URLs when useful."
|
| 241 |
+
),
|
| 242 |
+
},
|
| 243 |
+
{"role": "user", "content": search_prompt},
|
| 244 |
+
]
|
| 245 |
+
forwarded["temperature"] = 0
|
| 246 |
+
forwarded["max_tokens"] = min(int(payload.get("max_tokens", 160) or 160), 220)
|
| 247 |
+
result = proxy_json(base_url, "/v1/chat/completions", forwarded)
|
| 248 |
+
return str(result["choices"][0]["message"]["content"]).strip()
|
| 249 |
+
|
| 250 |
+
|
| 251 |
+
def make_handler(base_url: str, enable_search: bool, search_results: int) -> type[BaseHTTPRequestHandler]:
|
| 252 |
+
class Handler(BaseHTTPRequestHandler):
|
| 253 |
+
def do_GET(self) -> None:
|
| 254 |
+
if self.path == "/health":
|
| 255 |
+
json_response(self, 200, {"status": "ok"})
|
| 256 |
+
else:
|
| 257 |
+
json_response(self, 404, {"error": "not found"})
|
| 258 |
+
|
| 259 |
+
def do_POST(self) -> None:
|
| 260 |
+
length = int(self.headers.get("Content-Length", "0"))
|
| 261 |
+
try:
|
| 262 |
+
payload = json.loads(self.rfile.read(length).decode("utf-8"))
|
| 263 |
+
except json.JSONDecodeError:
|
| 264 |
+
json_response(self, 400, {"error": "invalid json"})
|
| 265 |
+
return
|
| 266 |
+
|
| 267 |
+
if self.path != "/v1/chat/completions":
|
| 268 |
+
try:
|
| 269 |
+
json_response(self, 200, proxy_json(base_url, self.path, payload))
|
| 270 |
+
except Exception as exc:
|
| 271 |
+
json_response(self, 502, {"error": str(exc)})
|
| 272 |
+
return
|
| 273 |
+
|
| 274 |
+
prompt = last_user_message(payload)
|
| 275 |
+
model = str(payload.get("model", "aurora-one"))
|
| 276 |
+
arithmetic = maybe_answer_arithmetic(prompt)
|
| 277 |
+
if arithmetic is not None:
|
| 278 |
+
json_response(self, 200, completion_payload(model, arithmetic))
|
| 279 |
+
return
|
| 280 |
+
deterministic = maybe_answer_deterministic(prompt)
|
| 281 |
+
if deterministic is not None:
|
| 282 |
+
json_response(self, 200, completion_payload(model, deterministic))
|
| 283 |
+
return
|
| 284 |
+
if enable_search and SEARCH_RE.search(prompt) and has_explicit_live_tool(payload):
|
| 285 |
+
try:
|
| 286 |
+
live_answer = maybe_answer_live_provider(prompt)
|
| 287 |
+
if live_answer is not None:
|
| 288 |
+
json_response(self, 200, completion_payload(model, live_answer))
|
| 289 |
+
return
|
| 290 |
+
answer = answer_from_search(base_url, payload, prompt, search_results)
|
| 291 |
+
json_response(self, 200, completion_payload(model, answer))
|
| 292 |
+
except Exception as exc:
|
| 293 |
+
json_response(self, 200, completion_payload(model, f"{UNKNOWN_FALLBACK} Search error: {exc}"))
|
| 294 |
+
return
|
| 295 |
+
if LIVE_RE.search(prompt) and not has_explicit_live_tool(payload):
|
| 296 |
+
json_response(self, 200, completion_payload(model, UNKNOWN_FALLBACK))
|
| 297 |
+
return
|
| 298 |
+
try:
|
| 299 |
+
json_response(self, 200, proxy_json(base_url, self.path, payload))
|
| 300 |
+
except Exception as exc:
|
| 301 |
+
json_response(self, 502, {"error": str(exc)})
|
| 302 |
+
|
| 303 |
+
def log_message(self, fmt: str, *args: Any) -> None:
|
| 304 |
+
print(f"{self.address_string()} - {fmt % args}")
|
| 305 |
+
|
| 306 |
+
return Handler
|
| 307 |
+
|
| 308 |
+
|
| 309 |
+
def main() -> None:
|
| 310 |
+
parser = argparse.ArgumentParser(description="OpenAI-compatible Aurora adapter in front of LM Studio.")
|
| 311 |
+
parser.add_argument("--listen-host", default="127.0.0.1")
|
| 312 |
+
parser.add_argument("--listen-port", type=int, default=8088)
|
| 313 |
+
parser.add_argument("--lmstudio-url", default="http://127.0.0.1:1234")
|
| 314 |
+
parser.add_argument("--enable-search", action="store_true")
|
| 315 |
+
parser.add_argument("--search-results", type=int, default=3)
|
| 316 |
+
args = parser.parse_args()
|
| 317 |
+
server = ThreadingHTTPServer(
|
| 318 |
+
(args.listen_host, args.listen_port),
|
| 319 |
+
make_handler(args.lmstudio_url, args.enable_search, args.search_results),
|
| 320 |
+
)
|
| 321 |
+
print(f"Aurora adapter listening on http://{args.listen_host}:{args.listen_port}")
|
| 322 |
+
print(f"Forwarding model calls to {args.lmstudio_url}")
|
| 323 |
+
print(f"Search enabled: {args.enable_search}")
|
| 324 |
+
server.serve_forever()
|
| 325 |
+
|
| 326 |
+
|
| 327 |
+
if __name__ == "__main__":
|
| 328 |
+
main()
|