Activity Feed

AI & ML interests

None defined yet.

Recent Activity

Zezoq8  updated a collection about 14 hours ago
Agent Collaborations
Zezoq8  updated a collection about 14 hours ago
Agent Collaborations
Zezoq8  updated a collection about 15 hours ago
Agent Collaborations
View all activity

Articles

DedeProGames 
posted an update 2 days ago
view post
Post
5643
how is this possible
  • 13 replies
·
DedeProGames 
posted an update 5 days ago
view post
Post
6255
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

▶ Play: DedeProGames/SLM-Tetris-Arena
📊 Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
·
DedeProGames 
posted an update 8 days ago
view post
Post
2872
🧱 SLM Tetris Arena: can a small language model play Tetris without ever being trained on it?

I built an arena where tiny decoder-only LMs (50K–250M params) play Tetris zero-shot. There is no fine-tuning and no game data. They only use what they picked up from pre-training on text.

How it works:
- For every piece, the engine simulates each legal placement and describes the result in plain English ("clears one line, creates no new holes, keeps the stack low…").
- The model never sees the grid. It reads each description, and the arena compares log P(" good move") with log P(" bad move"). The best-rated placement is played.
- Every player gets the same piece sequence, so it's a fair race.
- There are two protocols: Guided (the rules are in the prompt) and Blind (no rules, only pre-training knowledge).

Two ways to play:
- Match: pick any models (even your own, custom architectures welcome) and watch them play side by side on retro 8-bit boards.
- Ranked: press Play and the arena picks up to 4 models at random from a curated pool of 29. Nobody chooses their opponents, so Elo can't be farmed. Matches run on the server and count even if you close the tab.

First results (~225 ranked matches):
- gpt2 (124M) leads with 1283 Elo, but SupraNeo-4M (4M) is right behind at 1239. Next come LowOnMind-5M and BananaMind-2.1-Pico (1.5M!).
- Model size barely predicts Elo (r ≈ 0.06). Survival does (r ≈ 0.9): the models that avoid holes and keep the stack low are the ones that win.

Every ranked match (seed, model commit SHAs, scores, Elo before/after) is logged in a public dataset.

▶ Play: DedeProGames/SLM-Tetris-Arena
📊 Results: DedeProGames/lm-tetris-arena-results

Want your model in the Ranked pool? Drop it in the comments!
  • 1 reply
·
DedeProGames 
posted an update 9 days ago
view post
Post
150
🚀 Training GPT-U-20M on 2.6B tokens
DedeProGames 
posted an update 10 days ago
view post
Post
95
🚀 Possible new model in the GRM-3.2 family — GRM-3.2-Mist

A new model may be joining the GRM-3.2 family, originating from an experimental finetune currently under evaluation.

If this model passes internal testing and demonstrates strong performance, it will be released as GRM-3.2-Mist — a portable 1.7B model targeting reasoning and agentic coding tasks.

The model already exists and is currently undergoing evaluation. Follow this page for further updates:

OrionLLM

Question

10
#1 opened about 1 month ago by
GGUFGuy