Spaces:
Sleeping
Sleeping
Upload folder using huggingface_hub
Browse files- POST.md +35 -0
- profiler.py +5 -3
- requirements.txt +7 -3
POST.md
ADDED
|
@@ -0,0 +1,35 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# 30-in-15 · #12 — CSV → Insights Analyst (build-in-public post)
|
| 2 |
+
|
| 3 |
+
**Attach:** a screen-recording / screenshot of a CSV going in → charts + findings coming out.
|
| 4 |
+
**Demo:** https://insights.gritai.solutions · **Repo:** https://github.com/GritAI-Labs/csv-insights
|
| 5 |
+
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
📊 Project #12 of my "30 AI Projects in 15 Days": upload a CSV, get charts and a plain-English briefing back.
|
| 9 |
+
|
| 10 |
+
The twist is what it *won't* do. Most "AI analyzes your data" tools let the model read the numbers and write about them — which means the model can also make numbers up. This one can't. A deterministic pandas pass computes every statistic and every chart first; the AI only gets that finished profile and turns it into prose. If a figure is in the writeup, pandas computed it. The model literally never touches the arithmetic.
|
| 11 |
+
|
| 12 |
+
I went a step further on honesty. Small models love to bolt on units the data never had ("16.4°C", "3 mm") and invent methodology ("three standard deviations from the mean" when I used a different method). So there's a deterministic pass that strips guessed units, and the model is told to describe what the numbers show, not how they were computed. Percentages survive — because we compute those.
|
| 13 |
+
|
| 14 |
+
What you get: shape, missing-data and outlier flags, correlations, distributions, a trend line if there's a date column, top categories — and a short briefing that leads with what's interesting and ends with the questions the data could answer next.
|
| 15 |
+
|
| 16 |
+
Two tiers, same as the rest: this hosted demo writes the narrative with a hosted model; the GritAI Studio version runs the whole thing on your own local GPU fleet — free per-analysis, fully private, nothing leaves your network.
|
| 17 |
+
|
| 18 |
+
▶ https://insights.gritai.solutions
|
| 19 |
+
⭐ https://github.com/GritAI-Labs/csv-insights
|
| 20 |
+
|
| 21 |
+
---
|
| 22 |
+
|
| 23 |
+
## X / Twitter version (≤280)
|
| 24 |
+
|
| 25 |
+
📊 #12 of my 30 AI projects in 15 days: upload a CSV → charts + a plain-English briefing.
|
| 26 |
+
|
| 27 |
+
The catch: the AI never touches the numbers. pandas computes every stat + chart; the model only writes the story. Guessed units get stripped. Grounded by construction.
|
| 28 |
+
|
| 29 |
+
→ https://github.com/GritAI-Labs/csv-insights
|
| 30 |
+
|
| 31 |
+
---
|
| 32 |
+
|
| 33 |
+
*(Post to LinkedIn + X. After posting, append the permalinks below.)*
|
| 34 |
+
|
| 35 |
+
**Permalinks:** X: _____ · LinkedIn: _____
|
profiler.py
CHANGED
|
@@ -39,14 +39,16 @@ def _detect_datetime(s: pd.Series) -> pd.Series | None:
|
|
| 39 |
if s.dtype == object:
|
| 40 |
name = str(s.name).lower()
|
| 41 |
looks_datey = any(k in name for k in ("date", "time", "day", "month", "year", "timestamp"))
|
| 42 |
-
sample = s.dropna().head(50)
|
| 43 |
if sample.empty:
|
| 44 |
return None
|
| 45 |
with warnings.catch_warnings():
|
| 46 |
warnings.simplefilter("ignore") # dateutil "could not infer format" is expected here
|
| 47 |
-
|
|
|
|
|
|
|
| 48 |
if parsed.notna().mean() >= (0.6 if looks_datey else 0.95):
|
| 49 |
-
return pd.to_datetime(s, errors="coerce")
|
| 50 |
return None
|
| 51 |
|
| 52 |
|
|
|
|
| 39 |
if s.dtype == object:
|
| 40 |
name = str(s.name).lower()
|
| 41 |
looks_datey = any(k in name for k in ("date", "time", "day", "month", "year", "timestamp"))
|
| 42 |
+
sample = s.dropna().astype(str).head(50)
|
| 43 |
if sample.empty:
|
| 44 |
return None
|
| 45 |
with warnings.catch_warnings():
|
| 46 |
warnings.simplefilter("ignore") # dateutil "could not infer format" is expected here
|
| 47 |
+
# format="mixed" parses each value on its own terms — robust across pandas versions,
|
| 48 |
+
# where bare inference could coerce valid ISO dates to NaT and mis-type the column.
|
| 49 |
+
parsed = pd.to_datetime(sample, errors="coerce", format="mixed")
|
| 50 |
if parsed.notna().mean() >= (0.6 if looks_datey else 0.95):
|
| 51 |
+
return pd.to_datetime(s, errors="coerce", format="mixed")
|
| 52 |
return None
|
| 53 |
|
| 54 |
|
requirements.txt
CHANGED
|
@@ -1,5 +1,9 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
# profiler/charts/narrator use only the above + stdlib (urllib for the local Ollama backend).
|
| 5 |
# Optional Claude backend needs no extra package (raw HTTPS via urllib).
|
|
|
|
| 1 |
+
# Pinned to the exact versions verified locally (date detection + all 7 charts working).
|
| 2 |
+
# Unpinned (>=) let HF's build resolve different versions where date parsing behaved differently
|
| 3 |
+
# (date column mis-typed as categorical → trend chart dropped). Pin = reproduce the tested env.
|
| 4 |
+
flask==3.1.3
|
| 5 |
+
pandas==2.3.3
|
| 6 |
+
numpy==2.2.6
|
| 7 |
+
matplotlib==3.10.9
|
| 8 |
# profiler/charts/narrator use only the above + stdlib (urllib for the local Ollama backend).
|
| 9 |
# Optional Claude backend needs no extra package (raw HTTPS via urllib).
|