Unitos

Got a video, an audio file, an article, a research paper, a legal document, or a PDF assignment?Put it in Unitos Notebook.

Your understanding, your pace.

New here? Start your first project

Already have an account? Sign in

or

Sign-in creates your account and keeps your projects yours.

No billing information

Enter your email. Start your work now.

Two months of Unitos Premium, free. We never ask for a card to begin.

Contents Attention Is All You NeedNeural Machine Translation…ExtractNotes

Attention Is All You Need

Abstract

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.AssistantTell the assistant what to do with this selection…What does dropping recurrence and convolutions actually change?RunSimplifyVisualize UltraCommentAdd to notesLink across texts

Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU.

On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature.

1 Introduction

Recurrent neural networks, long short-term memory and gated recurrent neural networks in particular, have been firmly established as state of the art approaches in sequence modeling and transduction problems such as language modeling and machine translation.

Attention mechanisms have become an integral part of compelling sequence modeling and transduction models in various tasks, allowing modeling of dependencies without regard to their distance in the input or output sequences.

Simplifying…Simplified✕

We built a new, simpler model and named it the Transformer.

It relates the words of a sentence using attention only — each word looks at all the others. The two older building blocks, recurrence (reading word after word) and convolutions (sliding windows), are dropped completely.

Continue in a conversationPress a sentence for its original
Visualizing…VisualizationOpen✕
Recurrenceword after wordConvolutiona sliding windowAttention onlyevery word, every other, at once

“Dispensing with recurrence and convolutions entirely”: the two mechanisms crossed out, the one the Transformer keeps.

Assistanton the selection · ¶ 1✕

What does dropping recurrence and convolutions actually change?

No recurrence: positions are computed in parallel, so a training step no longer waits on the previous word. §3.2, Table 1

No convolutions: two distant words meet in one attention step, not through a stack of layers. §4

What replaces them: multi-head self-attention, feed-forward layers, and positional encodings to carry word order. §3.1, §3.5

Reply…Send
Read the selection aloudAnnotations · 3
NotesComparing 3 notesSide by sideStackedAdd note… ▾
Introduction✕
#a3f2···

Attention replaces recurrence

Every position attends to every other in one step — no sequential chain, so training parallelizes across the sentence.

“dispensing with recurrence and convolutions entirely” ¶ 1

2 sources · edited today

Training✕
#c71e···

Why it trains faster

Parallel attention means the whole batch runs at once: 12 hours on eight P100s to a new BLEU record.

“trained for 12 hours on eight P100 GPUs” ¶ 4

1 source · Mira, yesterday

Results✕
#9d40···

28.4 BLEU, and 41.8

Two records with one architecture. The En–Fr number came from 3.5 days of training.

“a new single-model state-of-the-art BLEU score of 41.8” ¶ 3

1 source · edited today

Introduction#a3f2✕

Attention replaces recurrence

Every position attends to every other in one step — no sequential chain, so training parallelizes across the sentence.

Why it trains faster

Parallel attention means the whole batch runs at once: 12 hours on eight P100s to a new BLEU record.

“dispensing with recurrence and convolutions entirely” ¶ 1 · “trained for 12 hours on eight P100 GPUs” ¶ 4

Training#c71e✕

Why it trains faster

Parallel attention means the whole batch runs at once: 12 hours on eight P100s to a new BLEU record.

“trained for 12 hours on eight P100 GPUs” ¶ 4

Results#9d40✕

28.4 BLEU, and 41.8

Two records with one architecture. The En–Fr number came from 3.5 days of training.

“a new single-model state-of-the-art BLEU score of 41.8” ¶ 3

Related work#4b8c✕

Where the idea came from

Self-attention was already in reading comprehension and entailment models; the Transformer is the first to rely on it alone.

“the first transduction model relying entirely on self-attention” ¶ 9

Open questions#e10d✕

Does the quadratic cost bite?

Attention is O(n²) in sequence length — fine at 512 tokens, ask Stitch what Scaling Laws says past that.

Merged #c71e into #a3f2Undo

Attention Is All You Need

The dominant sequence transduction models are based on complex recurrent or convolutional neural networks. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.1

Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.

Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU.

Recurrent models factor computation along the symbol positions of the input and output sequences, generating a sequence of hidden states ht as a function of the previous hidden state and the input for position t.

This inherently sequential nature precludes parallelization within training examples, which becomes critical at longer sequence lengths, as memory constraints limit batching across examples.

Attention mechanisms have become an integral part of compelling sequence modeling and transduction models in various tasks, allowing modeling of dependencies without regard to their distance in the input or output sequences. In all but a few cases, however, such attention mechanisms are used in conjunction with a recurrent network.

In this work we propose the Transformer, a model architecture eschewing recurrence and instead relying entirely on an attention mechanism to draw global dependencies between input and output. The Transformer allows for significantly more parallelization and can reach a new state of the art in translation quality after being trained for as little as twelve hours on eight P100 GPUs.

2 Background

The goal of reducing sequential computation also forms the foundation of the Extended Neural GPU, ByteNet and ConvS2S, all of which use convolutional neural networks as basic building block, computing hidden representations in parallel for all input and output positions.

In these models, the number of operations required to relate signals from two arbitrary input or output positions grows in the distance between positions, linearly for ConvS2S and logarithmically for ByteNet. This makes it more difficult to learn dependencies between distant positions.

⋮⋮Highlight¶ 1 · ···

“…based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.”

Core claim — everything else follows from it.

Search notes
Pending · 1← → to decide

The BLEU gain of 2 over ensembles is the headline result.

“improving over the existing best results, including ensembles, by over 2 BLEU” ¶ 2

AcceptRejectEnter · Backspace
Abstract 3+ Add note🎙

Attention replaces recurrence

Every position attends to every other in one step…

Why it trains faster

Parallel attention runs the whole batch at once…

Highlight“…dispensing with recurrence…”↗

28.4 BLEU, and 41.8

Two records with one architecture…

Introduction 1
Added to #c71e
#a3f2AbstractWrap text✕

Attention replaces recurrence

Every position attends to every other in one step — no sequential chain, so training parallelizes across the sentence.

“dispensing with recurrence and convolutions entirely” ¶ 1

Attention replaces recurrence
Every position attends to every other in one step — no sequential chain, so training parallelizes across the sentence. Contrast ¶ 4: the RNN builds one hidden state per step, so nothing runs in parallel.
DoneCancelSaved
Graph6 documents · 11 links
Attention Is All You Need
BERT
Lecture 12 · transcript
Scaling Laws
Course notes · web
Seminar audio

Lecture 12 ↔ Scaling Laws · 5 links

14:02 “compute grows with the square” ↔ ¶ 12

31:40 “the loss curve bends” ↔ Fig. 1

+ 3 more

Recommended links · 3

Attention ¶ 3 ↔ Scaling Laws ¶ 12Accept
BERT ¶ 2 ↔ Lecture 12 · 14:02Accept
Lecture 12 · 31:40 ↔ Seminar audio 08:15Accept

Generated content · 1

Attention cost across sequence lengthFrom the command: Connect the passages…Open
StitchThe assistant reads every document whole; a video or audio document as its transcript. It draws links between passages, or writes a new page from them, or both.✕
Every document (6)3 documents pickedAttention Is All You Need ✕BERT ✕Lecture 12 ✕Pick documents
Gather every paragraph that mentions…Connect the passages that answer…Find where these documents contradict each other

Connect the passages that answer how attention cost grows with sequence length

Reading the documents…■ Stop

All three answer it: the paper gives the quadratic term, BERT measures it at 512 tokens, the lecture explains why.

3 links proposed. Accept them under Recommended links.

Page written: Attention cost across sequence length Open

Read 3 of 3 documents · Attention Is All You Need 42 blocks read · BERT 61 blocks read · Lecture 12 312 transcript lines read

Connect the passages that answer how attention cost grows with sequence lengthWhat should the assistant do across these documents?Send
Attention Is All You NeedMJRHistoryShare

3 Model architecture

Most competitive neural sequence transduction models have an encoder-decoder structure. Here, the encoder maps an input sequence of symbol representations to a sequence of continuous representations.

Given z, the decoder then generates an output sequence of symbols one element at a time. At each step the model is auto-regressive, consuming the previously generated symbols as additional input.J

The Transformer follows this overall architecture using stacked self-attention and point-wise, fully connected layers for both the encoder and decoder, shown in the left and right halves of Figure 1, respectively.

3.1 Encoder and decoder stacks

The encoder is composed of a stack of N = 6 identical layers. Each layer has two sub-layers: a multi-head self-attention mechanism and a position-wise fully connected feed-forward network.

Jon
MJRMira owner · Jon editor · Rui viewerManage
Edits · 1 pendingJJon · 13:06 · ¶ 2

consuming the previously generated symbols conditioning on every symbol generated so far

AcceptRejectReply
MMiraSep 20, 13:02#e91b

Encoder = continuous representations

The encoder's job is only the mapping to z; generation is the decoder's. Worth a diagram for the group.

“maps an input sequence of symbol representations” ¶ 1

J
JonSep 20, 13:04Resolve×

Highlighted the sentence — anchor the diagram on it.

M
MiraSep 20, 13:05Resolve×

Agreed. Your edit to ¶ 2 reads better — accepting.

1 resolvedReply…Reply
HighlightJon · ¶ 1···

“the encoder maps an input sequence of symbol representations…”

R
RuiSep 19, 21:48Resolve

Link this to the BERT encoder paragraph? Same mapping.

History · Jon edited ¶ 2 · Mira added #e91b · Rui highlighted ¶ 13 today

Select a passage and the popover opens under it: the assistant’s command box, then Simplify, Visualize, Comment, Add to notes, Link across texts — the highlight colors right above. Simplify rewrites it beside the article, Visualize (Ultra) draws it, and the assistant answers a command on it in a chat card — every one saved under Annotations.

Break everything down, see all your progress.

Try it now for 2 months free!

Plans

Two tiers. Pick one and pay with Stripe.

Cancel any time.

2 Months Free Now

Scroll

Unitos Premium

$19.99 / month · $16.00 / month billed yearly

2 Months Free Now

Only necessary tools

Five tools. Each reads the documents and answers anchored to the passage it came from.

Simplify

The field weakens with distance. Its strength falls as the inverse square of the separation.

AssistantSimplifyColorsComment
SimplifiedTwice as far, a quarter as strong.

Rewrites the selection in plain words in a bubble beside the article. Press a sentence to light up the original it restates.

Extract

Where is flux defined?↵
“the flux through a surface is the field summed over it” ¶ 4Defines the term directly.
“the same flux, whatever the surface” ¶ 7Uses it three paragraphs on.

Ask the article one question. It scans the whole document and opens the extract page: your question on top, under it the quotes that answer it, each with a caption saying how.

Assistant

The field weakens with distance. Its strength falls as the inverse square of the separation.

AssistantSimplifyColorsComment
Why a quarter?Distance is squared ¶ 2 — the passage's own step, stated again at ¶ 5.

Type or speak a question or a command about the selection. A question is answered from the passages across the article that match it, each cited with a ¶ chip that jumps to it.

Stitch

Drop documents here
3 documents · one project
PDFPrincipia, Book III
VideoLecture 4 — Inverse square
WebGauss on flux (web)
StitchFind where these documents contradict each other

Add two or more documents to a project. Stitch, the assistant across them, gathers passages, draws links between documents, finds contradictions, or writes a new page from all of them.

Graph view

ListGraph
3 documents · 5 linksPDFVideoWeb
3 links between these documents

The project as a graph: each node a document, each curve the links between two of them. A thicker line means more links; dashed means a recommended link awaiting acceptance. Drag to arrange, click a node to open it.

Perfect your workflow with notes

Notes sit beside the article, quote it, absorb everything you find, and open full screen to compare.

A miniature Google Doc for every note

Headings, lists, checklists, quotes, bold, italic, underline, four text colors, indent, pictures up to 25 MB — every tool on the bar, every one with a typed shortcut. Saves as you go.

NoteSaving…Saved
H1H2H3•–1.☐❝BIU⇤⇥
Inverse square, in one page
The field falls with the square of distance.
•Twice as far → a quarter as strong
•Flux through any closed surface: unchanged
Check against Principia, Book III

Merge with AI

2 selectedMerge with AIPin
Distance
Field falls with r². Twice as far, a quarter as strong.
Flux
Same flux whatever the surface, the passage says.
Merging with AI…
Merged noteInverse square and flux
•Field falls with r² ¶ 2
•Flux unchanged by the surface ¶ 5
2 notes merged into oneUndo

Select two notes with the circle at their top right and press Merge with AI. It writes the one note that takes both notes’ place: the key points of each, each with the quotes that support it. Undo puts them back.

Quotes straight from the reader

The flux through a surface is the field summed over it, and it holds whatever the surface.

“flux through a surface is the field summed over it”
FluxWhere the article defines it:
“flux through a surface is the field summed over it”¶ 4
Let go to put the quote here

Drag a passage out of the article and let it go inside a note. A caret shows where it lands; the quote arrives on its own line with a ¶ chip that jumps back to its exact words.

Annotations become notes

Annotations
a quarter as strong
Check the 1687 derivation
flux is unchanged
Jump
Note · floatingOpen questionsDoes the 1687 result need both claims?
a quarter as strong
Check the 1687 derivation
Drop to merge this annotation in

While a note floats over the article, drag any highlight or comment by its grip onto the floating card. It merges into the note, source and all.

A note absorbs anything

Drag in a highlight, a comment, a quote, an extracted answer, or a picture up to 25 MB. Each one lands in the note with its source attached, and jumps back to the passage it came from.

AnnotationExtractionCommentQuote ¶ 4Picture · 25 MB
NoteInverse square
Highlights · commentsQuotes · extractionsPictures
5 sources attached

Compare, then merge

Select notes on the notes full page and press Compare: they open in one screen, side by side or stacked. Read them against each other — then merge the two that say the same thing.

Side by sideStackedMerge with AI
Compare · 3 notesSide by sideStacked
NewtonForce falls with r²; the 1687 proof rests on it.“the inverse square”
GaussFlux through a closed surface is fixed by what it encloses.“whatever the surface”
Lecture 4Both say one thing: geometry sets the fall-off.— transcript, 12:40
NewtonForce falls with r²; the 1687 proof rests on it.“the inverse square”
Lecture 4Both say one thing: geometry sets the fall-off.— transcript, 12:40
2 selectedMerge with AI
Merging with AI…
Merged noteGeometry sets the fall-off
•Force falls with r² — Newton’s 1687 proof rests on it ¶ 2
•The lecture puts both under one idea 12:40
2 notes merged into oneUndo

Read and write together

One project, every collaborator on it: the same article, the same notes, every change signed.

Share the project

BLShare
Share this project with collaborators.
lena@lab.eduEditor ▾Add
BBobby
bobby@lab.edu
Owner
LLena
lena@lab.edu
Editor ▾

The owner adds collaborators by email as Editor or Viewer, the Google Docs way. Everyone who is here now shows at the top of the reader; a viewer reads, an editor writes.

Notes, together

Pending · 1⏎ accept · ⌫ reject
LLena· Pending

Gauss gives the same fall-off from geometry alone — note it next to Newton.

AcceptReject
BBobby· Sep 16

Inverse square: twice as far, a quarter as strong.

LLena Sep 17, 10:12
Does this hold inside the shell too?
Resolve

Every note carries who wrote it. A collaborator’s note lands as pending in your section: accept it with Enter, reject with Backspace. Reply under any note; replies resolve and collapse behind a count.

Annotate the same reader

The field weakens with distance. Its strength falls as the inverse square of the separation, and the flux through any closed surface is unchanged.

BL
LLenacommentedJump

This is Gauss, not Newton — worth a link to the other document.

BAgreed — linking it now.Reply

Highlights and comments from everyone sit on one article, each with its author’s badge. Click the comment icon beside the text to open the thread, and jump from any annotation to its passage.

Comment on each other’s edits

EditsEdited text shows in color
EditLSep 17, 10:20
Was

falls with the distance

Now

falls as the inverse square of the distance

Revert
BBobby 10:24
Cite ¶ 2 for this, or it reads as ours.
LLena 10:26
Done — chip added.
Resolve

Edits to the article show in the Edits tab, newest first, each signed and each with what it was and what it is now. Reply right under an edit, revert it, or resolve the thread once it is settled.

See the whole history

HistoryBLM
Who changed what, and when.
LLena edited text
falls as the inverse square of the distance
10:20
BBobby merged notes
Inverse square and flux
09:52
MMei removed a paragraph
Duplicate statement of the law
09:31
LLena added a link
→ Gauss on flux (web)
09:14

Every edit, merge and deletion in the project, newest first, signed by the account that did it. Press a person’s badge to see only their work; restore any note to an earlier version.

Plan the work together

AssistantDocumentProject
ContradictionsGaps
Reading 14 notes…
Contradiction

B says the fall-off is Newton’s; L says it follows from geometry alone. Decide which the notes lead with.

Open question

Nobody has covered the field inside the shell. Assign

In the side panel the Assistant works at project scope for everyone in it: ask a question across all the notes, or run a check — Contradictions finds notes that disagree, Gaps finds thin sections, claims without support and open questions to take on next.

Every document kind

Imported whole and as it is.

PDFWeb pageImageVideoAudio

Wait, there's more…

Everything in Premium, and the night sky above it

Unitos Ultra

$29.99 / month · $30.00 / month billed yearly

1Visualize

Everybody loves visuals.

Get to the bottom of complex analysis and descriptions by visualizing it. Select text; Ultra returns a diagram, a drawing, or a short animation.

Ultra only

Visualizedrawing the passage…
PassageAnchorDerivationPicture
2Tool conversations

Keep asking.

Explain+, Simplify+, Analyze+, Visualize+ — any tool’s answer becomes a chat, same passage, same anchor, as long as you need.

Ultra only

Simplify +Explain +Analyze +Visualize +
SimplifiedTwice as far means a quarter as strong.
Then what holds at the boundary?
The flux through the surface, unchanged — the passage’s own claim, three paragraphs on. ¶ 5
3Works offline

Write on a plane.

Every edit saves on your device and syncs, in order, the moment you reconnect.

Four changes, queuedSynced
✓A note edited on the plane✓A passage highlighted, offline✓A section renamed, waiting✓An image dropped into a note
4Runs and annotations

Twice the runs. No cap.

2× the Extract runs of Premium, and unlimited highlights, comments and links.

PremiumExtract
Ultra2× the runs
∞Annotations — highlights, comments, links. No cap.
5Storage

Five times the room.

5× Premium’s space for your documents, images and video.

Premiumstandard
Ultra5× the space
DocumentsImagesVideo

Only the necessary and valuable.

More on the way…

Choose Now

Two tiers. Pick one and pay with Stripe.

Cancel any time.

Unitos Premium

$19.99 / month

or $16.00 / month billed yearly · save 19%

2 Months Free Now
  • Reading, notes, anchoring, export
  • Documents: PDF, web page, image, video, audio
  • AI: derivations, assistant, extract, glossary, conversion
  • Sharing and collaboration
  • Extract: a set number of runs per document
  • Annotations up to the included allowance · standard storage · images up to 25 MB in notes

Unitos Ultra

$29.99 / month

or $30.00 / month billed yearly

  • Everything in Unitos Premium
  • Visualize: the selection as a picture — a diagram, a drawing, or a short animation
  • Tool conversations: continue a Simplify, Analyze, or Visualize card into a conversation
  • Offline copies: a project's pages and images kept in the browser, so it opens without a network
  • 2× the Extract runs, and unlimited annotations
  • 5× the storage space