Thompson sampling
Heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem
Nº Q7795822 ★★
Uncommon · Knowledge
Thompson sampling
Heuristic for choosing actions that addresses the exploration-exploitation dilemma in the multi-armed bandit problem
Thompson sampling, named after William Rae Thompson, is a heuristic for choosing actions that address the exploration–exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.
Last price
—
Floor price
—
7-day median
—
30-day sales
0
30-day range
—
In circulation
0
Price history
median
low – high
sales
No sales in this period
Show table
| Date | median | Low | High | sales |
|---|
Sales history
- Last sale
- —
- 30-day average
- —
- 30-day low
- —
- 30-day high
- —
- Sales 7d
- 0
- Sales 30d
- 0
No sales yet.
Anonymous sales: no buyer or seller shown. Figures count player-to-player sales only.
From Wikipedia
Thompson sampling, named after William Rae Thompson, is a heuristic for choosing actions that address the exploration–exploitation dilemma in the multi-armed bandit problem. It consists of choosing the action that maximizes the expected reward with respect to a randomly drawn belief.
Text: Wikipédia, CC BY-SA 4.0. · Image: Nguiard (CC BY 4.0) ·