These commands only read information — they change nothing. Run the
block for your OS, then use the table to read your own numbers
against the ideas above.
What you should see Format-Hex prints the file's bytes as a table of hexadecimal values with an ASCII column on the right, so the exact byte values and the readable text both appear side by side.
Note Windows has no built-in file type-detector like the Unix file command, so there is no detected-type line here; Format-Hex shows the raw bytes plus the ASCII column instead.
What you should see file reports a detected type such as ASCII text for /etc/hosts; od -c prints the bytes as characters, showing the file is plain readable text with newline bytes shown as escapes.
Note macOS ships the BSD versions of file and od, so wording and column spacing differ slightly from the GNU tools on Linux.
What you should see file inspects the contents of /etc/hosts and reports a detected type such as ASCII text; od -c prints each byte as its printable character or a backslash escape, so the readable host entries appear directly and control bytes such as newlines show as escapes like \n.
Where each idea shows up in your output
Concept
Windows
macOS
Linux
Detected file type
—
file /etc/hosts
file /etc/hosts
Raw bytes of the file
Format-Hex C:\Windows\System32\drivers\etc\hosts
od -c /etc/hosts
od -c /etc/hosts
The same ideas, as prose
These are the exact fragments the model serves — also available as an
ordered study guide.
Bytes mean nothing until you agree what they mean
A file on disk, or a message crossing a network, is never anything more than a run of bytes. The bytes do not announce whether they are a photo, a spreadsheet, or a paragraph of English. Meaning is not stored alongside them; it is supplied by whoever reads them, according to a rule they already know.
A data format is exactly that rule: the shared agreement about how to turn a sequence of bytes back into the numbers, characters, and structures someone meant to record. The writer encodes meaning into bytes by one convention and the reader decodes it by the same one like a language both the writer and the reader must share, or the same words land as nonsense. Pick the wrong convention and the same bytes come back as nonsense, which is why a format mismatch is so often the reason a file simply will not open.
So a format is a contract between two programs that may never meet, written at different times by different people. Everything that follows is about the kinds of contracts in common use: whether the bytes are meant to be read by a human or only by a machine, which byte stands for which letter, and how a format carries not just values but the shape those values are arranged in.
Readable, or compact
The value one thousand stored as text takes four readable character bytes, one per digit, while stored as native binary it takes only two bytes.
Every format falls on one side of a basic split. A text format encodes data as human-readable characters, so the bytes spell out something a person can open in any editor and read. A binary format packs the same data into machine-native bytes chosen for the machine's convenience, not the reader's, so opening it in an editor shows garbage.
The split is a genuine trade, not a ranking. Writing the number one thousand as the four characters 1000 costs one byte per digit and is legible anywhere; the same value in a machine's native form can fit in two bytes and be read without any digit-by-digit conversion. Text buys you readability and portability across systems that disagree about almost everything else. Binary buys you smaller size and faster reading and writing, at the cost of needing a program to make any sense of it.
Neither wins outright. A config file you edit by hand wants to be text; a stream of measurements written millions of times a second wants to be binary. The useful instinct is to know which one you are looking at, because reaching for a text tool on a binary file is a common and confusing waste of time.
Which byte is which letter
Calling something a text format only pushes the question down a level: which byte stands for which character? That mapping is a character encoding, and text is readable only when the reader uses the same one the writer did. Get it wrong and the letters come back scrambled, accented characters most of all.
ASCII is the old, small answer. It is a 7-bit code, which gives it exactly 128 characters: the English letters, the digits, common punctuation, and a handful of control codes. That was enough for early English-language computing and nothing more, so it has no room for the rest of the world's writing.
UTF-8 is the answer that won. It encodes every character in Unicode, the universal catalogue of characters, using a variable number of bytes: one byte for the simplest characters and up to four for the rarest. Its decisive trick is that it is byte-for-byte compatible with ASCII for those first 128 characters, so every plain ASCII file is already valid UTF-8 with nothing to convert. That backward compatibility is much of why UTF-8 became the default encoding of the web and of most modern text.
Rows, columns, and commas
In CSV, accented commas are the delimiters that split each line into fields, while a comma inside a double-quoted field is part of the value and does not split it.
The simplest way to store a table as text is to write one record per line and separate the fields within a line by a chosen mark. CSV (comma-separated values) does exactly that: fields divided by commas, records divided by line breaks like a spreadsheet flattened into lines of text, with a chosen mark standing in for the lines between the columns. It is about as simple as a format gets, which is why almost every spreadsheet and database can read and write it.
The simplicity has a sharp edge. If the delimiter is a comma, what happens when a field's own value contains a comma, like a name written last-name-first? The convention (set out in RFC 4180) is to wrap any field containing a comma, a double quote, or a line break in double quotes, and to escape a double quote inside such a field by doubling it. So a value with an internal comma travels safely as a single quoted field instead of splitting into two.
That one rule is the whole reason a naive split-on-every-comma parser eventually mangles someone's data, and the reason it is usually wiser to reach for a real CSV library than to write the parsing yourself. The format is trivial until a value contains the very character that structures the file.
Nesting, named and typed
A small JSON object parses into a tree: the root object has keyed edges to a number, a nested array, and a nested object that in turn holds a boolean.
A flat table of rows and columns cannot easily hold a value that is itself a list, or a record tucked inside another record. Structured text formats can. JSON (JavaScript Object Notation) is the one you will meet most: readable text that represents values nested to any depth, with each value carrying its own type.
JSON offers a small, fixed set of types. There are two that hold other values: an object, a set of named fields, and an array, an ordered list. And there are the values that stand on their own: a string, a number, a boolean written true or false, and null for the deliberate absence of a value. Because an object's field can itself be another object or array, the shape and the names travel with the data like values tucked into labeled boxes within boxes, so each one's name and place travel with it, so a reader learns what each value is called and how the pieces fit without a separate description.
That self-describing, nested quality is why JSON dominates configuration files and the messages passed between web services. It is still text, so a person can read and edit it, but unlike a flat delimited file it can carry the full shape of a structure, not just a grid.
Saying what shape the data must take
A format tells you how to read bytes into values, but not whether those values are the ones you were expecting. A schema fills that gap. It is a declaration of the shape data must take: which fields must be present, what type each one is, and sometimes what range or pattern a value is allowed to fall in. With a schema in hand, a program can check a document against it before doing anything else.
That check is validation, and it earns its keep for two reasons at once. The first is correctness: catching a missing field or a number where a string belongs, right at the boundary, instead of letting it cause a baffling failure deep inside later. The second is security. Data arriving from outside your program is untrusted by default, and validating it against an expected shape is a front-line defense against malformed or hostile input, which is one of the oldest and most reliable sources of bugs and breaches. Data that fails the schema is rejected before it is ever trusted.
The discipline is simple to state: decide what well-formed input looks like, write it down as a schema, and refuse anything that does not match. What you skip validating is exactly what you are trusting blindly, and untrusted input is precisely the input you cannot afford to trust.
Turning a structure into bytes and back
A running program holds data as structures in memory, but memory does not outlive the program and cannot be sent down a wire. To save a structure to a file or ship it to another machine, it has to become a flat run of bytes. Serialization is that outbound step: writing an in-memory structure out as a sequence of bytes. Parsing, or deserialization, is the inbound step: reading those bytes back into a structure a program can use.
Every format in this concept is one choice of how to do that pair. JSON and CSV serialize to text, which keeps the result readable. A binary serialization format makes the opposite choice, packing the structure into machine-native bytes instead of characters. The result is smaller on disk and faster to write and read, because there are no digits to spell out and no characters to scan, but you give up being able to read it without the matching program.
So the trade from the readable-versus-compact split shows up here in concrete form: reach for a text serialization when a human needs to read or edit the result, and for a binary one when size and speed matter more than legibility. What both share is that serialization and parsing must agree on the format exactly, because bytes written by one convention and read by another come back as nonsense.
Choosing the right shape for the job
Choosing a data format is a design decision with real consequences, and the trades in this concept point straight at the answer. When a service needs to send a nested, self-describing message and people will read it while debugging, JSON earns its place. When the data is a plain table moving between tools that all speak it, CSV is the low-friction choice. When there is a great deal of data and size or speed dominates, a compact binary format pays for its unreadability. There is no single best format, only the one that fits the job.
The other load-bearing skill is humbler: being able to look at raw bytes when a file will not parse. A format mismatch, a wrong character encoding, or an unexpected binary blob all look the same from the outside, which is a file that should open and does not. Dropping down to inspect the actual bytes, and asking the system what type it thinks the file is, turns a mystery into a fact.
That is why these fundamentals stay load-bearing well past this concept. The larger systems built on top of them, from data stores to the pipelines that feed modern AI tools, all still come down to bytes agreed upon on both ends. Get the format right and everything above it has a solid floor to stand on.